Hi, Thank you for the review.
We found that sk_forward_alloc is a plain int updated by a non-atomic RMW (sk_forward_alloc_add(), where even the read side is not READ_ONCE), and its writers span three unrelated lock domains: - socket lock, process context: SO_RESERVE_MEM and TX grants (__sk_mem_schedule(), sk_forced_mem_schedule()); - receive-queue lock, softirq: UDP RX charges and the udp_rmem_release() fold; - no lock at all: sk_mem_charge()/sk_mem_uncharge() from skb_set_owner_r() and skb destructors, and the sk_mem_reclaim() fold itself. Any cross-domain pair loses an update. A lost charge is still returned in full by the matching skb free, so it resurfaces as a phantom surplus in sk_forward_alloc, and the next fold hands it back to the memcg via __sk_mem_reduce_allocated() -> mem_cgroup_sk_uncharge() - an uncharge with no matching charge. We first tried to fix it with locks, and hit three walls: - the socket lock cannot be used: its holders free skbs (e.g. tcp_recvmsg()), and the destructor's sk_mem_uncharge() would need to re-acquire it - recursion. That is why these helpers are lockless in the first place; - a new per-socket spinlock serializes the RMWs but not the bug: __sk_mem_schedule() publishes the grant before the memcg charge, and the charge may sleep (GFP_KERNEL, memcg reclaim/OOM), so no spinlock can cover both steps; a fold in that window can still refund pages whose charge afterwards fails; - such a lock would also sit on the per-packet charge/uncharge paths, exactly the hot path the cacheline layout around sk_forward_alloc was tuned to keep cheap. Are there any good solutions to solve this problem? On 9/24/2026 4:26 PM, Eric Dumazet wrote:
On Thu, Sep 24, 2026 at 9:36 AM Cai Xinchen <[email protected]> wrote:The memcg socket accounting currently charges pages to the memory cgroup per grant (__sk_mem_schedule() publishing forward allocation) and refunds them later from skb destructors. The refund side folds per-skb "was this charged" snapshots back into the socket balance under concurrent lockless RMW, and races there can drive the memcg socket balance negative, ending with: page_counter underflow WARNING: ... mm/page_counter.c ... page_counter_cancel() This series flips the model: a socket is charged its whole memory budget (sk_sndbuf + sk_rcvbuf + sk_reserved_mem) to its memcg when the budget is established or grows, and refunded when the budget shrinks or the socket dies. Grants and per-skb charge/uncharge stop touching the memcg entirely, so the racy refund pairing has no code left to go wrong: refunds can never exceed charges and the balance cannot underflow by construction. Tested: full arm64 build with 0 warnings; each intermediate state compiles (bisectable); tools/testing/selftests/cgroup builds clean. Runtime validation on the workload that used to trigger the underflow is pending.Charging sk_sndbuf + sk_rcvbuf upfront to memory.current is not viable, especially for servers handling large numbers of connections (e.g. 1 million TCP sockets): We specifically went in the exact opposite direction in commit 4890b686f408 ("net: keep sk->sk_forward_alloc as small as possible") to make sure idle sockets hold zero forward-allocated memory and non-idle sockets hold less than one page (4 KB) in sk_forward_alloc. 1. Massive phantom memory charges and false OOMs: With default sysctl_tcp_wmem[1] (16 KB) and sysctl_tcp_rmem[1] (128 KB), 1 million completely idle TCP sockets immediately charge 144 GB to the cgroup's memory.current while holding 0 bytes of actual packet buffers. Worse, once TCP autotuning grows sk_rcvbuf / sk_sndbuf during a short burst (up to tcp_rmem[2] = 6 MB and tcp_wmem[2] = 4 MB by default), TCP does not shrink sk_rcvbuf or sk_sndbuf when the queues drain and the connection becomes idle again. 1 million long-lived, mostly-idle connections would permanently pin hundreds of GBs (up to several TBs) of non-existent memory in memory.current, forcing the memcg into constant reclaim thrashing of real page cache/anon pages and triggering premature memcg OOM kills. 2. Overcommitted caps vs. physical reservations: sk_sndbuf and sk_rcvbuf are per-socket upper bounds that are heavily overcommitted across sockets, not reservations (unlike SO_RESERVE_MEM). Comparing this to vm_committed_as is flawed: vm_committed_as tracks virtual address space overcommit globally and is never charged to memcg's memory.current for the exact same reason. 3. Broken memcg limit enforcement (memory.max bypass): __sk_mem_raise_allocated() drops mem_cgroup_sk_charge() completely, while sk_memcg_budget_sync() ignores charge failures ("the new budget is used uncharged"). When a cgroup reaches memory.max, sk_memcg_budget_sync() fails in sock_init_data_uid(), tcp_init_sock(), setsockopt(SO_SNDBUF/SO_RCVBUF), or autotuning, yet sk_sndbuf and sk_rcvbuf are still raised. Subsequent skb allocations in __sk_mem_schedule() will then allocate real physical memory without charging the memcg at all. 4. Unnecessary struct sock bloat and hot-path overhead: - Adds 8 bytes (sk_memcg_budget + sk_memcg_budget_lock) to struct sock (even when !CONFIG_MEMCG). - Acquires spin_lock_bh(&sk->sk_memcg_budget_lock) inside sk_mem_reclaim(). - Every TCP socket creation charges rmem_default + wmem_default (416 KB) in sock_init_data_uid() and immediately uncharges 272 KB in tcp_init_sock(). If you are hitting a page_counter underflow race in socket memcg accounting, please share the exact race / stack trace and fix the underlying accounting bug rather than charging uncommitted buffer limits.

