From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) (using TLSv1 with cipher DHE-RSA-AES256-SHA (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 843F2C9830E for ; Thu, 24 Sep 2026 09:25:51 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 9BAB76B008A; Thu, 24 Sep 2026 05:25:50 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 991FE6B008C; Thu, 24 Sep 2026 05:25:50 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 8A8BE6B0092; Thu, 24 Sep 2026 05:25:50 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0010.hostedemail.com [216.40.44.10]) by kanga.kvack.org (Postfix) with ESMTP id 5BBC96B008A for ; Thu, 24 Sep 2026 05:25:50 -0400 (EDT) Received: from smtpin10.hostedemail.com (lb01a-stub [10.200.18.249]) by unirelay10.hostedemail.com (Postfix) with ESMTP id CC25EC02AF for ; Thu, 24 Sep 2026 09:25:49 +0000 (UTC) X-FDA: 85248123618.10.982C499 Received: from canpmsgout12.his.huawei.com (canpmsgout12.his.huawei.com [113.46.200.227]) by imf01.hostedemail.com (Postfix) with ESMTP id 9B6FD4000C for ; Thu, 24 Sep 2026 09:25:46 +0000 (UTC) Authentication-Results: imf01.hostedemail.com; dkim=pass header.d=huawei.com header.s=dkim header.b=oujNsmgU; spf=pass (imf01.hostedemail.com: domain of caixinchen1@huawei.com designates 113.46.200.227 as permitted sender) smtp.mailfrom=caixinchen1@huawei.com; dmarc=pass (policy=quarantine) header.from=huawei.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1790241947; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=bDikpWMchPRZnYe6ZTjPxkVDUAVMs72pAdd9atbppWw=; b=oQUYWUor/kma7XZ7yB1M3ZU5nC4NLh7rO9tslfvebYcN0yvKxTZDTjQsoO4Bp91fr8gnus mSZ7YBJoaAQ6lg2sovFiVrzkUFoFS1VeylxvSzlaaLlvQG6ZyziZp1bbO6YZ5lAGxt56U3 dlpPLVWTtvnMVtnobLRbcDsrqgYrmbg= ARC-Authentication-Results: i=1; imf01.hostedemail.com; dkim=pass header.d=huawei.com header.s=dkim header.b=oujNsmgU; spf=pass (imf01.hostedemail.com: domain of caixinchen1@huawei.com designates 113.46.200.227 as permitted sender) smtp.mailfrom=caixinchen1@huawei.com; dmarc=pass (policy=quarantine) header.from=huawei.com ARC-Seal: i=1; a=rsa-sha256; d=hostedemail.com; s=arc-20220608; cv=none; t=1790241947; b=xpD2PyHEASvC5SyHCm+R29LE+nOrD2bmg1R+hEsQKcHit0nZdV1TpDpGngFdVatYroTnps SJmoUOhaM+FwFnBsFf9O5e1O/+P3VeoQ5zWslaP8axrRsf6O9v2H8Ih4IzaWfTe5MDW6Sh SXv8wYBbWbjSDaeJwmIJ5Y7ARQXqF5s= dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=bDikpWMchPRZnYe6ZTjPxkVDUAVMs72pAdd9atbppWw=; b=oujNsmgUBwdI/eT1H+Qs5F7p2O9MDX6RLH5o2hJQlaLXAb/j6LdmuAxKAUikKwjx4ELt2b2y2 lgUdHwxOESQ7bGcylb6iraoXxxUeyVASQYTIWGMFuLFmCa9f6T1UfMS9oTgO2W7hua9cXwHXMQz TcR+D/VUrVG60YNnOWDC2RI= Received: from mail.maildlp.com (unknown [172.19.162.92]) by canpmsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hr7Rh5MhQznTVd; Thu, 24 Sep 2026 17:13:36 +0800 (CST) Received: from whupemk100010.china.huawei.com (unknown [7.152.184.41]) by mail.maildlp.com (Postfix) with ESMTPS id 91E8840586; Thu, 24 Sep 2026 17:25:40 +0800 (CST) Received: from [10.67.109.91] (10.67.109.91) by whupemk100010.china.huawei.com (7.152.184.41) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.45; Thu, 24 Sep 2026 17:25:36 +0800 Message-ID: Date: Thu, 24 Sep 2026 17:25:35 +0800 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH RFC -next 0/5] net: charge socket memory budget to memcg upfront To: Eric Dumazet CC: , , , , , , , , , , , , , , , , , , , , , , , , , , , References: <20260924080219.1036588-1-caixinchen1@huawei.com> Content-Language: en-US From: Cai Xinchen In-Reply-To: Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 8bit X-Originating-IP: [10.67.109.91] X-ClientProxiedBy: kwepems100002.china.huawei.com (7.221.188.206) To whupemk100010.china.huawei.com (7.152.184.41) X-Rspamd-Server: rspam05 X-Rspamd-Queue-Id: 9B6FD4000C X-Stat-Signature: ng3zgqeayrkaxsudiqyicf1pt1ayeg55 X-Rspam-User: X-HE-Tag: 1790241946-85193 X-HE-Meta: U2FsdGVkX19BOEDAqo5hQhUsLRN24p/mbnmkFTCXCgJfwegNRAxZ9LmDfwNb9OfYg5D2ib/Y/K86mxCs4WMnBt6k/uT4vUmqZfNjf4pr9M8Mgm/+nlD5NLbcTTRjbv486x3v5HayRd4iDQw05NGqyj7HGqPR19lp4uxyZZo80YK4sHgweRtJinMZp1ds8fbzs3lh64ZY3FmnJtTRdlO2HWBSnLyNU/G/5T5EiS13dEAnTQlTYK5JcZVuMevpcTFTQvP7Pk7OyNPzicN8t/vlowQZgk3aOYEVfl3SgQjgClOSk1QnRHXyhXwWI8cs9VulHrKVF+oDJK3r0Su4/319Tpr3vrcHDXrlHn+lP3ui2cjLzlwIlkFKkQvfuURnnbFgR0O+vvsOO3ixG2v24kn0BFh+z849UWwgBYzvVUpHQiS0yU3KHVlsG6BR99UMMOfaxcpUm2KX5rgSnxoQ2LqlngYYmqeJl32X87hPVxO7KIXgttDFGnjHa/zynSzfNczphOhlwhjOutc8mhoNmfVtF5/sXNZ24Ilgq+jGc/RC8AF8UplqxeYGGkl1GL/d306lS69evSTVri6d/KhZyfnWKAjGc6TmNB5LGIkPkilAaodwgAxq3MAsLqJTKAjS2dkbANdFkdi2Js1Hcsl7WjnpEAo3wMP5AV9bADYz0dp1Dqr+t0yx/zlm3H4ccqQygrfZPDDzvPBk2M4yuhxkku5VWRl9tKRdlDNplpNu5FhtkMct2ren6F6RdqL0aoIqX1Gs/MIdx6voJXF+R4Mx8IeREu/XiLUP/FoEBRZ1gExIfPsOoNUhzvUVFWHRvy348989VgtxhzOr6lq3mwZbfdbKBMh6eN9QjEi2kdVbhLj0hi/0dyg3Ymt+NYW71lrZVREy/2iKxR7IH3p/EofF244JUG1NWnLd/FVoW0u6M/mkOiLT0zj2Z26ihWp9TVbJDqFiJpqfd4qkRhprUvCscgQ LPsHxl9Y /KTZUSJvAehfHuup+aUBQoYkhWmno6sAfR2ENOa8otbpn2ZR9AoWbWezV/EYiO+0gGSNl/r/JRLBfhANkgmuzhePWwOYmlPzx1TjTDcTDvScyFS8CAdSCnrqgumZnHvEQZ1MNFJg3Ad+rUMqxxk8PklOS6izGK+OEo8WWfQldsXK6bfFwOmelidQbyzhaIUP/FE7QThLMXez2Y2Ispy3qFBVXL36Vm4psQxYhw2pdXddnrzwV9HUQzjFNxwDr0zD/Z/a8aPTK93Zquk7rKvRS8KJTFwKEzm8c4K5pMODPjbYJn1DGZsbHpXQ4HuySvsQzU8xJvPV8pFzR9sUrEYJVa6vXxFG2CRV1r6O0C1J3ZYp8AcSIwkX/mgO3GA== Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: Hi, Thank you for the review. We found that sk_forward_alloc is a plain int updated by a non-atomic RMW (sk_forward_alloc_add(), where even the read side is not READ_ONCE), and its writers span three unrelated lock domains:   - socket lock, process context: SO_RESERVE_MEM and TX grants      (__sk_mem_schedule(), sk_forced_mem_schedule());   - receive-queue lock, softirq: UDP RX charges and the      udp_rmem_release() fold;   - no lock at all: sk_mem_charge()/sk_mem_uncharge() from      skb_set_owner_r() and skb destructors, and the sk_mem_reclaim()      fold itself. Any cross-domain pair loses an update.  A lost charge is still returned in full by the matching skb free, so it resurfaces as a phantom surplus in sk_forward_alloc, and the next fold hands it back to the memcg via __sk_mem_reduce_allocated() -> mem_cgroup_sk_uncharge() - an uncharge with no matching charge. We first tried to fix it with locks, and hit three walls:   - the socket lock cannot be used: its holders free skbs (e.g.      tcp_recvmsg()), and the destructor's sk_mem_uncharge() would      need to re-acquire it - recursion.  That is why these helpers      are lockless in the first place;   - a new per-socket spinlock serializes the RMWs but not the bug:      __sk_mem_schedule() publishes the grant before the memcg charge,      and the charge may sleep (GFP_KERNEL, memcg reclaim/OOM), so no      spinlock can cover both steps; a fold in that window can still      refund pages whose charge afterwards fails;   - such a lock would also sit on the per-packet charge/uncharge      paths, exactly the hot path the cacheline layout around      sk_forward_alloc was tuned to keep cheap. Are there any good solutions to solve this problem? On 9/24/2026 4:26 PM, Eric Dumazet wrote: > On Thu, Sep 24, 2026 at 9:36 AM Cai Xinchen wrote: >> The memcg socket accounting currently charges pages to the memory >> cgroup per grant (__sk_mem_schedule() publishing forward allocation) >> and refunds them later from skb destructors. The refund side folds >> per-skb "was this charged" snapshots back into the socket balance >> under concurrent lockless RMW, and races there can drive the memcg >> socket balance negative, ending with: >> >> page_counter underflow >> WARNING: ... mm/page_counter.c ... page_counter_cancel() >> >> This series flips the model: a socket is charged its whole memory >> budget (sk_sndbuf + sk_rcvbuf + sk_reserved_mem) to its memcg when the >> budget is established or grows, and refunded when the budget shrinks >> or the socket dies. Grants and per-skb charge/uncharge stop touching >> the memcg entirely, so the racy refund pairing has no code left to go >> wrong: refunds can never exceed charges and the balance cannot >> underflow by construction. >> >> Tested: full arm64 build with 0 warnings; each intermediate state >> compiles (bisectable); tools/testing/selftests/cgroup builds clean. >> Runtime validation on the workload that used to trigger the underflow >> is pending. >> > > Charging sk_sndbuf + sk_rcvbuf upfront to memory.current is not > viable, especially for servers handling large numbers of connections > (e.g. 1 million TCP sockets): > > We specifically went in the exact opposite direction in commit > 4890b686f408 ("net: keep sk->sk_forward_alloc as small as possible") > to make sure idle sockets hold zero forward-allocated memory and > non-idle sockets hold less than one page (4 KB) in sk_forward_alloc. > > > 1. Massive phantom memory charges and false OOMs: > With default sysctl_tcp_wmem[1] (16 KB) and sysctl_tcp_rmem[1] > (128 KB), 1 million completely idle TCP sockets immediately charge > 144 GB to the cgroup's memory.current while holding 0 bytes of > actual packet buffers. > Worse, once TCP autotuning grows sk_rcvbuf / sk_sndbuf during a > short burst (up to tcp_rmem[2] = 6 MB and tcp_wmem[2] = 4 MB by > default), TCP does not shrink sk_rcvbuf or sk_sndbuf when the queues > drain and the connection becomes idle again. 1 million long-lived, > mostly-idle connections would permanently pin hundreds of GBs (up to > several TBs) of non-existent memory in memory.current, forcing the > memcg into constant reclaim thrashing of real page cache/anon pages > and triggering premature memcg OOM kills. > > 2. Overcommitted caps vs. physical reservations: > sk_sndbuf and sk_rcvbuf are per-socket upper bounds that are heavily > overcommitted across sockets, not reservations (unlike SO_RESERVE_MEM). > Comparing this to vm_committed_as is flawed: vm_committed_as tracks > virtual address space overcommit globally and is never charged to > memcg's memory.current for the exact same reason. > > 3. Broken memcg limit enforcement (memory.max bypass): > __sk_mem_raise_allocated() drops mem_cgroup_sk_charge() completely, > while sk_memcg_budget_sync() ignores charge failures ("the new budget > is used uncharged"). When a cgroup reaches memory.max, > sk_memcg_budget_sync() fails in sock_init_data_uid(), tcp_init_sock(), > setsockopt(SO_SNDBUF/SO_RCVBUF), or autotuning, yet sk_sndbuf and > sk_rcvbuf are still raised. Subsequent skb allocations in > __sk_mem_schedule() will then allocate real physical memory without > charging the memcg at all. > > 4. Unnecessary struct sock bloat and hot-path overhead: > - Adds 8 bytes (sk_memcg_budget + sk_memcg_budget_lock) to struct sock > (even when !CONFIG_MEMCG). > - Acquires spin_lock_bh(&sk->sk_memcg_budget_lock) inside > sk_mem_reclaim(). > - Every TCP socket creation charges rmem_default + wmem_default > (416 KB) in sock_init_data_uid() and immediately uncharges 272 KB > in tcp_init_sock(). > > If you are hitting a page_counter underflow race in socket memcg > accounting, please share the exact race / stack trace and fix the > underlying accounting bug rather than charging uncommitted buffer limits.