Confirmed - i915 is the cause. Booted with modprobe.blacklist=i915, everything else identical (cursor_blink=1, vtcon0 bound, same guest):
vmwgfx COMMAND stays 0007 (I/O+ Mem+), no "VGA decodes changed" in dmesg 0 blocks / 600 s inline cmdbuf pool: 1 in use, 4 capacity, 1 page Normal boot, for comparison: vmwgfx COMMAND 0004 (I/O- Mem-, all three BARs [disabled]) 2931 blocks / 600 s = 4.885/sec inline cmdbuf pool: 66041 in use, 66044 capacity, 16511 pages Four measurements, same 600 s window and method, cursor_blink=1 throughout: i915 loaded, COMMAND 0004 (as i915 leaves it) : 4.885/sec i915 rmmod'd after boot, COMMAND 0004 : 4.882/sec i915 loaded, COMMAND forced to 0007 via setpci : 4.882/sec i915 blacklisted from boot, COMMAND 0007 : 0.000/sec The third line is the one I did not expect: restoring the decode bits at runtime does not stop the leak, and neither does unloading i915. So the damage is not simply "decode is off right now" - once vmwgfx has been through the arbiter interaction it stays broken even after I/O+Mem come back. Only never clearing them works. That rules out setpci as a workaround, and is consistent with cc6ed470caa2 being the right fix since it prevents the clearing rather than repairing it. Full logs, both boots: https://www.google.com/url?q=https://gist.github.com/oGi4i/c7c4ffcb9726df7bba4bca5835fbfd66&source=gmail&ust=1786219041539000&sa=E Guest: Debian 13, kernel 6.12.101+deb13-amd64, ESXi 8.0.3 build-25205. i915 is a passed-through CoffeeLake UHD 630 (8086:3e92). cc6ed470caa2 is not in 6.12.y - by my reading it first appears in v7.1. Ville, would you consider a backport to 6.12.y? Happy to test a patch here; the reproducer is continuous and I can toggle it with cursor_blink. On Fri, Aug 07, 2026 07:16 AM, Zack Rusin <[email protected]> wrote: > On Thu, Aug 6, 2026 at 12:08 AM Zack Rusin <[email protected]> > wrote: > > > > Sorry, are you using llm for replies? The email starts with multiple > > paragraphs just to say, "Oh, yeah,", which, as a human, makes it > > difficult to communicate :) If you could provide the data/logs and let > > us do the analysis then it's just going to go a lot smoother. Can you > > upload your logs somewhere? They haven't gone through. If you could > > upload your vmware.log and the guest system log somewhere (some > > pastebin is fine) that would be great. > > > > z > > > > On Wed, Aug 5, 2026 at 5:29 PM Дима Дьячков <[email protected]> wrote: > > > > > > On Wed, Aug 5, 2026 at 10:36 PM Zack Rusin <[email protected]> > wrote: > > > > > > Thanks both. Answers inline, but one thing up front: > > > > > > **The pools request revised the severity. The leak is ~5x larger than > > > I reported.** Each leaked command buffer holds a 1024-byte DMA pool > > > block in addition to the 256-byte kmalloc header, and DMA pool memory > > > is not accounted in SUnreclaim - so every measurement in my original > > > report was blind to 80% of it. Corrected figure is ~516 MiB/day, not > > > ~103 MiB/day. Detail below. > > > > > > > Could you send us the full system log? (ideally journalctl -b, if not > > > > then, dmesg, or /var/log/kern.lorg). I suspect you have some command > > > > buffer errors in the kernel log that are causing this. > > > > > > No command buffer errors - the log is clean. Full journalctl -b and > > > dmesg are attached/available on request, but the summary is: > > > > > > - 19 vmwgfx lines total, all at boot (t=3.84-3.91s), all > > > informational. Nothing from vmwgfx after that at all. > > > - Zero vmwgfx errors, warnings, or *ERROR* lines. > > > - The only DRM error anywhere is "i915 *ERROR* PPS state mismatch", > > > from a passed-through Intel GPU used for video transcoding. Not > > > vmwgfx, and present from boot regardless of the leak. > > > - 5 call traces in dmesg, all of them cgroup OOM kills of a > > > kube-state-metrics pod hitting its own memory limit. Unrelated > > > (and since fixed). > > > > > > One boot line is relevant though, since it confirms the allocator: > > > > > > [drm] Using command buffers with DMA pool. > > > > > > So the pool is being used as intended - it is only the release side > > > that never happens, silently. > > > > > > > Based on your analysis disabling the cursor_blink is probably a > > > > better workaround, i.e. "echo 0 > > /sys/class/graphics/fbcon/cursor_blink" > > > > > > Confirmed, and it identifies the exact driver. Measured over 600s > > > windows by watching the vmw_cmdbuf_alloc live-object count in > > > /sys/kernel/debug/slab/kmalloc-256/alloc_traces: > > > > > > cursor_blink=1 : 4.887 obj/sec (2932 objects / 600 s) > > > cursor_blink=0 : 0.003 obj/sec (2 objects / 600 s) > > > > > > A ~1600x reduction; effectively zero. > > > > > > The rate pins the mechanism precisely. fbcon's cursor timer runs at > > > HZ/5, and this kernel is CONFIG_HZ=250, so 50 jiffies = 200 ms = 5.0 > > > Hz. Measured 4.887 Hz. That is **one leaked command buffer per cursor > > > blink tick**, with the small shortfall being scheduler jitter. It also > > > agrees with the 4.89 obj/sec I derived independently from the age= > > > field in my first mail. > > > > > > So on a completely idle headless guest, the blinking cursor alone > > > leaks ~516 MiB/day. > > > > > > This is a much better workaround than the ones I listed originally - > > > it keeps the console fully usable rather than blanking it. Thank you. > > > > > > One caveat for anyone else applying it: after an unbind/rebind cycle > > > of vtcon0 I observed cursor_blink read back 0 while the leak was still > > > running at the full 4.887 obj/sec. The attribute appears to > > > under-report in that state, so it is worth confirming the leak > > > actually stopped rather than trusting the read-back. > > > > > > > Besides the system logs i'd be also intersted in pool info, i.e. cat > > > > /sys/bus/pci/devices/<vmwgfx card id>/pools . > > > > > > This is the interesting one. > > > > > > $ cat /sys/bus/pci/devices/0000:00:0f.0/pools > > > poolinfo - 0.1 > > > vmwgfx inline cmdbuf 53875 53876 1024 13469 > > > vmwgfx cmdbuf 0 0 64 0 > > > > > > Read as name / blocks-in-use / capacity / block-size / pages: > > > > > > - 53,875 of 53,876 blocks in use - the pool has grown to exactly > > > accommodate every allocation and sits pinned at ~100% utilisation. > > > - 13,469 pages = 52.6 MiB held. (13469 * 4096/1024 = 53876, which > > > matches the stated capacity, so the reading is self-consistent.) > > > - The non-inline "vmwgfx cmdbuf" pool is completely empty, so this > > > is specific to the inline command buffer path. > > > > > > Taken at the same moment: > > > > > > $ sudo awk '/vmw_cmdbuf_alloc/{print $1; exit}' \ > > > /sys/kernel/debug/slab/kmalloc-256/alloc_traces > > > 53872 > > > > > > 53,872 kmalloc-256 headers against 53,875 pool blocks - a difference > > > of 3, which is just allocations in flight between the two reads. So it > > > is 1:1: every leaked command buffer pins both a 256-byte header and a > > > 1024-byte DMA pool block, 1280 bytes total. > > > > > > Sanity check: 53,875 blocks / 4.887 per sec = 11,024 s = 3.06 hours of > > > active leaking, which matches the accumulated time the console was > > > bound on this boot. > > > > > > Since dma_pool memory comes from dma_alloc_coherent rather than the > > > slab allocator, none of it appears in SUnreclaim. That is why my > > > original 103 MiB/day figure was low by 5x - it only ever counted the > > > kmalloc headers. > > > > > > On Wed, Aug 5, 2026 at 8:30 PM Ian Forbes <[email protected]> > wrote: > > > > > > > > Please recompile with KMEMLEAK and/or KASAN and share any output. > > > > > > If the additional info provided is not enough, then I will recompile. > > > Two notes on how I plan to do it: > > > > > > Debian ships neither CONFIG_DEBUG_KMEMLEAK nor CONFIG_KASAN, so this > > > needs a custom build. I will do it in a clean throwaway Debian VM > > > rather than on the host I reported from, for two reasons: > > > > > > 1. That host runs ZFS via DKMS, so its kernel is tainted > > > (P/O/E). A report from an untainted kernel with nothing but > > > Debian base installed is worth more, and removes any question > > > about whether ZFS, Kubernetes or the passed-through i915 are > > > involved. > > > 2. KASAN's memory overhead is not viable on that host, but is fine > > > on a bare VM. > > > > > > I will send kmemleak output, and KASAN separately if it turns up > > > anything - though since kmemleak reports unreferenced allocations and > > > KASAN reports memory errors, I expect kmemleak is the relevant one > > > here unless you suspect corruption in the release path. > > > > > > Happy to test patches in the meantime; the reproducer is a stock > > > Debian VMware guest with a text console, and I can toggle the leak on > > > and off with cursor_blink. > > > > > > Thanks, Dmitry > > > > > > > > > ср, 5 авг. 2026 г. в 22:36, Zack Rusin <[email protected]>: > > > > > > > > On Wed, Aug 5, 2026 at 12:59 PM Дима Дьячков <[email protected]> > wrote: > > > > > > > > > > Hi, > > > > > > > > > > I have a reproducible, quantified kernel memory leak in vmwgfx on > a headless > > > > > VMware guest. It is attributed to a specific call site with > slab_debug, and I > > > > > have a clean causal test that stops it. I could not find this > reported anywhere > > > > > - I searched the dri-devel archives, the mainline and 6.12.y git > logs, and the > > > > > distro trackers. > > > > > > > > > > Summary > > > > > ======= > > > > > > > > > > On a headless VMware guest with the text console bound to the > vmwgfx > > > > > framebuffer, the fbdev-emulation damage worker issues a full > atomic commit for > > > > > every console damage event, and vmwgfx leaks a command buffer on > each one. The > > > > > result is unreclaimable kernel slab growing linearly at ~103 > MiB/day, never > > > > > reclaimed under memory pressure, released only by reboot. > > > > > > > > > > Environment > > > > > =========== > > > > > > > > > > Kernel: 6.12.100+deb13-amd64 (Debian 13 stock, tracks 6.12.y) > > > > > Guest: VMware Virtual Platform, BIOS 6.00; vmwgfx drives the > virtual display > > > > > Console: headless, administered over SSH; a single agetty on the > console > > > > > fb0: vmwgfxdrmfb > > > > > vtcon0: "(M) frame buffer device", bind=1 > > > > > Config: CONFIG_HZ=250, CONFIG_SLUB_DEBUG=y, CONFIG_STACKDEPOT=y > > > > > (CONFIG_DEBUG_KMEMLEAK is not set in Debian's kernel) > > > > > > > > > > i915 is also loaded on this host (a passed-through GPU used for > video > > > > > transcoding) but is unrelated - the console framebuffer is vmwgfx. > > > > > > > > > > Symptom > > > > > ======= > > > > > > > > > > SUnreclaim grows linearly at ~103 MiB/day and is never reclaimed > under pressure. > > > > > The growth is confined to the plain kmalloc-256 cache - notably > *not* > > > > > kmalloc-cg-256, so it is plain GFP_KERNEL from an in-kernel > allocator rather > > > > > than anything cgroup-accounted. > > > > > > > > > > Attribution > > > > > =========== > > > > > > > > > > Booted with slab_debug=U,kmalloc-256, then read > > > > > /sys/kernel/debug/slab/kmalloc-256/alloc_traces, which ranks by > live object > > > > > count: > > > > > > > > > > 41025 vmw_cmdbuf_alloc+0x40/0x200 [vmwgfx] waste=328200/8 > age=254700/1301825/2351179 pid=9-208926 cpus=0-5 > > > > > vmw_cmdbuf_alloc+0x40/0x200 [vmwgfx] > > > > > vmw_cmdbuf_reserve+0x142/0x170 [vmwgfx] > > > > > vmw_cmd_ctx_reserve+0x24/0x60 [vmwgfx] > > > > > vmw_stdu_bind_st+0x6b/0x100 [vmwgfx] > > > > > vmw_stdu_primary_plane_atomic_update+0xa8/0x1d0 [vmwgfx] > > > > > drm_atomic_helper_commit_planes+0x112/0x320 [drm_kms_helper] > > > > > drm_atomic_helper_commit_tail+0x48/0x90 [drm_kms_helper] > > > > > vmw_atomic_commit_tail+0x1b/0x80 [vmwgfx] > > > > > commit_tail+0x91/0x130 [drm_kms_helper] > > > > > drm_atomic_helper_commit+0x11a/0x140 [drm_kms_helper] > > > > > drm_atomic_commit+0xa9/0xe0 [drm] > > > > > drm_atomic_helper_dirtyfb+0x1c6/0x2a0 [drm_kms_helper] > > > > > drm_fbdev_ttm_helper_fb_dirty+0x216/0x310 [drm_ttm_helper] > > > > > drm_fb_helper_damage_work+0x94/0x170 [drm_kms_helper] > > > > > process_one_work+0x174/0x330 > > > > > worker_thread+0x191/0x2f0 > > > > > > > > > > The next-largest entry in that cache held 1056 objects - this one > is ~30x > > > > > everything else combined. > > > > > > > > > > Quantitative confirmation > > > > > ========================= > > > > > > > > > > With CONFIG_HZ=250, from the age=min/avg/max field above: > > > > > > > > > > - max age 2,351,179 jiffies = 9405 s = the full uptime at time of > reading > > > > > - avg age is almost exactly half of max, which is the signature of > uniform > > > > > allocation with zero frees > > > > > - 41,025 objects over the 8386 s during which allocation was active > > > > > = 4.89 objects/sec = 4.30 MiB/h = ~103 MiB/day > > > > > > > > > > That matches the independently measured SUnreclaim growth rate to > within noise. > > > > > > > > > > Causal test > > > > > =========== > > > > > > > > > > Unbinding the text console from the framebuffer stops it instantly: > > > > > > > > > > # echo 0 > /sys/class/vtconsole/vtcon0/bind > > > > > > > > > > After that, the minimum object age jumped from 26 jiffies to > 254,700 and kept > > > > > climbing - i.e. no new allocations at all. SUnreclaim also dropped > ~93 MiB > > > > > immediately as the fbdev shadow buffer was released, and has been > flat-to- > > > > > falling since. > > > > > > > > > > Two things I checked so you don't have to > > > > > ========================================= > > > > > > > > > > This is not a recent regression. > drivers/gpu/drm/vmwgfx/vmwgfx_cmdbuf.c has had > > > > > no functional change between v6.12 and current mainline - only the > command > > > > > buffer ID addition (922f9de) and the treewide kmalloc -> > kmalloc_obj > > > > > conversions. vmwgfx_stdu.c is similarly untouched in the relevant > paths, and > > > > > nothing in the 6.12.y stable series touches cmdbuf, damage or > dirtyfb handling. > > > > > So this looks long-standing rather than newly introduced. > > > > > > > > > > It is also not caused by the fbdev-ttm conversion, which is the > obvious suspect > > > > > given the stack. drm_fbdev_ttm was a pure rename of > drm_fbdev_generic > > > > > (aae4682e5d66, v6.10), and vmwgfx has used the generic DRM fbdev > emulation since > > > > > v6.2 (df42523c12f8, which deleted vmwgfx_fb.c). What did change at > v6.2 is that > > > > > console damage now goes through a full atomic commit, which the old > > > > > vmwgfx_fb.c did not do - but the leak itself is on the vmwgfx side. > > > > > > > > > > Possibly related > > > > > ================ > > > > > > > > > > There is an open, unresolved report of a different vmwgfx leak in > the same > > > > > atomic-commit path - Yuma Kakei, 12 Mar 2026, "[BUG] vmwgfx: > Severe Slab memory > > > > > leak (radix_tree_node) with Wayland Atomic KMS" (kernels 6.17 / > 7.0-rc3, KWin > > > > > Wayland with Atomic KMS, stops with KWIN_DRM_NO_AMS=1). Ian Forbes > proposed a > > > > > patch; the reporter confirmed on 15 Mar 2026 that it did not fix > it. > > > > > > > > > > I do *not* claim these are the same bug - different slab caches, > different > > > > > trigger, different kernels. But both are vmwgfx leaking on every > atomic commit > > > > > and stopping dead when the atomic path stops, so they may share a > root cause. > > > > > > > > > > Workarounds > > > > > =========== > > > > > > > > > > - echo 0 > /sys/class/vtconsole/vtcon0/bind > > > > > runtime, reversible, stops it immediately > > > > > > > > > > - drm_kms_helper.fbdev_emulation=0 on the kernel command line > > > > > cleaner for a headless host; disables fbdev emulation outright so > no > > > > > damage worker exists. Must be set at boot. > > > > > > > > > > - modprobe.blacklist=vmwgfx > > > > > heavy-handed; loses all console output > > > > > > > > > > All three cost the virtual console, which is fine headless but is > not a general > > > > > fix. > > > > > > > > > > Happy to test patches, run instrumented kernels, or provide > further data - the > > > > > system reproduces this continuously and I can rebind the console > at will to > > > > > toggle it on and off. > > > > > > > > > > Thanks, Dmitry > > > > > > > > Could you send us the full system log? (ideally journalctl -b, if not > > > > then, dmesg, or /var/log/kern.lorg). I suspect you have some command > > > > buffer errors in the kernel log that are causing this. Based on your > > > > analysis disabling the cursor_blink is probably a better workaround, > > > > i.e. "echo 0 > /sys/class/graphics/fbcon/cursor_blink" . Besides the > > > > system logs i'd be also intersted in pool info, i.e. cat > > > > /sys/bus/pci/devices/<vmwgfx card id>/pools . > > From the vmware.log that I was sent in private what is happening is > that when the i915 driver that you passthrough to the vm loads, it > takes the VGA arbiter, the abiter clears the io mem on vmwgfx and > never restores it. So vmwgfx is broken by the loaded i915 driver > because our communication channel to the device is wiped from > underneath us. We keep queueing cb's and can't free them because we're > waiting for some kind of notification from the device that it > processed them but that will never happen. > > I think upstream commit cc6ed470caa2 (“drm/i915/vga: Avoid VGA arbiter > during intel_vga_disable() for iGPUs”) > > https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=cc6ed470caa2 > will fix this. You can give it a try and if it fixes it for you, you > can ask Ville for a backport. > > z >
