On Thu, Sep 24, 2026 at 9:36 AM Cai Xinchen <[email protected]> wrote:
>
> The memcg socket accounting currently charges pages to the memory
> cgroup per grant (__sk_mem_schedule() publishing forward allocation)
> and refunds them later from skb destructors. The refund side folds
> per-skb "was this charged" snapshots back into the socket balance
> under concurrent lockless RMW, and races there can drive the memcg
> socket balance negative, ending with:
>
> page_counter underflow
> WARNING: ... mm/page_counter.c ... page_counter_cancel()
>
> This series flips the model: a socket is charged its whole memory
> budget (sk_sndbuf + sk_rcvbuf + sk_reserved_mem) to its memcg when the
> budget is established or grows, and refunded when the budget shrinks
> or the socket dies. Grants and per-skb charge/uncharge stop touching
> the memcg entirely, so the racy refund pairing has no code left to go
> wrong: refunds can never exceed charges and the balance cannot
> underflow by construction.
>
> Tested: full arm64 build with 0 warnings; each intermediate state
> compiles (bisectable); tools/testing/selftests/cgroup builds clean.
> Runtime validation on the workload that used to trigger the underflow
> is pending.
>
Charging sk_sndbuf + sk_rcvbuf upfront to memory.current is not
viable, especially for servers handling large numbers of connections
(e.g. 1 million TCP sockets):
We specifically went in the exact opposite direction in commit
4890b686f408 ("net: keep sk->sk_forward_alloc as small as possible")
to make sure idle sockets hold zero forward-allocated memory and
non-idle sockets hold less than one page (4 KB) in sk_forward_alloc.
1. Massive phantom memory charges and false OOMs:
With default sysctl_tcp_wmem[1] (16 KB) and sysctl_tcp_rmem[1]
(128 KB), 1 million completely idle TCP sockets immediately charge
144 GB to the cgroup's memory.current while holding 0 bytes of
actual packet buffers.
Worse, once TCP autotuning grows sk_rcvbuf / sk_sndbuf during a
short burst (up to tcp_rmem[2] = 6 MB and tcp_wmem[2] = 4 MB by
default), TCP does not shrink sk_rcvbuf or sk_sndbuf when the queues
drain and the connection becomes idle again. 1 million long-lived,
mostly-idle connections would permanently pin hundreds of GBs (up to
several TBs) of non-existent memory in memory.current, forcing the
memcg into constant reclaim thrashing of real page cache/anon pages
and triggering premature memcg OOM kills.
2. Overcommitted caps vs. physical reservations:
sk_sndbuf and sk_rcvbuf are per-socket upper bounds that are heavily
overcommitted across sockets, not reservations (unlike SO_RESERVE_MEM).
Comparing this to vm_committed_as is flawed: vm_committed_as tracks
virtual address space overcommit globally and is never charged to
memcg's memory.current for the exact same reason.
3. Broken memcg limit enforcement (memory.max bypass):
__sk_mem_raise_allocated() drops mem_cgroup_sk_charge() completely,
while sk_memcg_budget_sync() ignores charge failures ("the new budget
is used uncharged"). When a cgroup reaches memory.max,
sk_memcg_budget_sync() fails in sock_init_data_uid(), tcp_init_sock(),
setsockopt(SO_SNDBUF/SO_RCVBUF), or autotuning, yet sk_sndbuf and
sk_rcvbuf are still raised. Subsequent skb allocations in
__sk_mem_schedule() will then allocate real physical memory without
charging the memcg at all.
4. Unnecessary struct sock bloat and hot-path overhead:
- Adds 8 bytes (sk_memcg_budget + sk_memcg_budget_lock) to struct sock
(even when !CONFIG_MEMCG).
- Acquires spin_lock_bh(&sk->sk_memcg_budget_lock) inside
sk_mem_reclaim().
- Every TCP socket creation charges rmem_default + wmem_default
(416 KB) in sock_init_data_uid() and immediately uncharges 272 KB
in tcp_init_sock().
If you are hitting a page_counter underflow race in socket memcg
accounting, please share the exact race / stack trace and fix the
underlying accounting bug rather than charging uncommitted buffer limits.