David Woodhouse <[email protected]> writes: > On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote: >> David Woodhouse <[email protected]> writes: >> >> > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote: >> > > >> > > === What if my provider doesn't deal with pages? >> > > >> > > Frank and David Woodhouse [2] have use cases for providers that use >> > > PFNs, and this is definitely something a generic interface must >> > > support. >> > > >> > > Here are some options I can think of: >> > > >> > > 1. Don't define .alloc_folio(), instead define .alloc_pfn() >> > > 2. Refactor .alloc_folio() to .alloc_pfn() >> > > >> > > Either way, I think guest_memfd should still be the primary manager of >> > > the memory. >> > >> > Thanks for working on this. >> > >> >> Thanks for the quick reply! >> >> > I'm not sure what you mean by 'primary manager' here but my use case is >> >> Naming is hard! See comment on revocation below. >> >> > for the provider to be the ultimate arbiter of who owns which PFN, and >> > revocation of the same. Think of it like a filesystem backed by >> > external memory. Each guest's memory is a file, and a guest can >> > *donate* specific pages of its memory to other guests (which in the >> >> Putting donation aside first - I don't have a good picture of how a >> guest would tell the host that it's willing to share a page - can't >> comment. Would love to find out more. >> >> > context of the *interface* only means that we need revocation). > > That's implementation-specific. Think of things like Nitro Enclaves > where a guest 'donates' its memory back to the hypervisor to be used by > another microvm. But that interface isn't the guest_memfd concern; as I > said, in the context of the guest_memfd interface, there is only > revocation: that page went away, whatever the reason. >
Okay we're on the same page then. :) guest_memfd only needs to know that the page went away, guest_memfd cannot say no. >> Going with this revocation part, beginning with a more basic use case: >> I'm thinking that the provider can notify guest_memfd that the page or >> PFN is going away. > > Right. That part I have working in what I was posting. > [5] > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06 > That is similar to what I had in mind except the parameters, I was thinking inode, offset, size. kvm_gmem_invalidate_range() in your patch [5] does have a kvm_gmem prefix, but it's really a direct call to tell KVM MMU to zap the gfn range, guest_memfd doesn't get to update its own state. >> guest_memfd cannot say no, but it does have to tell KVM to zap the page >> from stage 2, and for CoCo it may need to do more stuff. >> >> The notification needs some pointer, so I was thinking that on .attach() >> the provider would also note down which guest_memfd instance the >> provider should notify. >> >> The notifier could tell guest_memfd that some (offset, size) is going >> away, and breaking down mappings in the EPT/NPT should be something >> guest_memfd tells KVM to do, I think. Are you thinking that the provider >> would directly break the mappings down in the EPT/NPT? > > I think telling KVM is the right thing to do. > Glad to hear that :) >> Does this notification mechanism work for you? IIUC regular DAX would >> need it too, like if the device is unplugged for example. > > Sure. As long as it works and efficiently covers everyone's use cases, > I'm not going to be overly opinionated a priori about *how* it works. > > All such bets are off once we see the code, of course :) > >> > Should I be trying to rework parts of my series to live on top of this? >> >> I hope to gather more feedback at LPC. It'd be nice to get early >> feedback if you think it can work for you! The overall design still >> needs feedback though, don't take this as the final direction. >> >> > Things I have working but which I don't see here include: >> > >> >> This sounds like 4 or more different features! I think they do work with >> this resource fd proposal. >> >> > • PFN support (which you mentioned). >> >> Yup, guest_memfd needs some way to track PFNs in addition to folios. I >> think for tracking we can reference DAX, the part I haven't looked at is >> how to let guest_memfd handle events. Like if there's memory failure on >> a PFN, how do we let guest_memfd handle it? > > Telling the guest about correctable and uncorrectable errors, you mean? > I would have thought that's up to the VMM, until the point where (see > revocation)? > Informing the guest is definitely the userspace VMM's action to take. guest_memfd's role here is updating it's own tracking, and perhaps splitting folios. When guest_memfd supports huge pages, one thing high on the wishlist is to split the huge folio and zap only the page where there was an error and not the huge page. Advantage of splitting: If the failure happened in the last 4K, the first 1G - 4K can continue to be used by the guest. If the guest never touches the last 4K, it doesn't have to know about the failure at all. >> guest_memfd needs to at least be notified, so that it can handle the KVM >> stuff: zapping and special stuff for CoCo. Whether the provider or >> guest_memfd gets to handle it first can be worked out :) >> >> > • IOMMUFD support via dma-buf. >> >> Is this about having guest_memfd work with dmabuf fds? > > It's about having the provider of memory (again, think of it as a daxfs > if you will) able to tell *both* KVM and IOMMU which pages are where — > and revoke them at will. > > The *implementation* involved dmabuf fds, and we already discussed > which wraps which, on the first posting of my series. But maybe we'll > conclude that the right answer is for IOMMUFD to recognise this > guest_memfd 'provider' as a first-class citizen, and we won't need to > masquerade as dmabuf? > >> Fuad also brought up that Android uses dmabuf as an interface for CMA >> memory. [3] IIUC the dmabuf fd is mmap()ed, then the userspace address >> is set up in a memslot for the guest. To make this CMA memory >> CoCo-friendly, we want to use it with guest_memfd. >> >> Jason response was that guest_memfd should probably just directly get >> memory from CMA [3]. >> >> I'm not super familiar with CMA or dmabufs, would either have a suitable >> "resource fd" to pass to guest_memfd like how HugeTLBfs has a mount fd? >> The "resource fd" should ideally have no way to mmap or read/write the >> memory directly. >> >> [3] >> https://lore.kernel.org/all/CA+EHjTxZ0N3Tfnid404B4tkb_E+Z8mODTHTgBPiF6=bwzp7...@mail.gmail.com/ > > I think where the memory comes *from* is an implementation detail for > the provider. It provides a PFN (or folio, if you must). All else is > not the business of KVM or IOMMUFD. > So we already have dmabuf --wraps--> CMA (and others) We could have [a] guest_memfd --wraps--> dmabuf --wraps--> CMA or [b] guest_memfd --wraps--> CMA I agree guest_memfd shouldn't care where the PFN comes from, but this does impact where the provider_ops code is implemented though. If we go with [b], CMA needs to find a non-dmabuf way for userspace to get some resource fd. If we go with [a], dmabuf folks need to support this :) >> Or is this about letting IOMMUFD get pages to map in IOMMU page tables >> via an fd+offset instead of via mmap()-ed userspace addresses as in >> IOMMU_IOAS_MAP_FILE, and that guest_memfd isn't a supported kind of fd >> (yet)? >> >> I think guest_memfd should support IOMMU_IOAS_MAP_FILE. One use case is >> Confidential IO [4], where SNP would need to have private memory >> (guest_memfd) mapped in the IO page tables. > > Right, I think that's what I meant with 'first class citizen' above? > I agree, yes, I would like guest_memfd to be first class alongside other memory providers. >> The missing parts here are that guest_memfd and iommufd need to >> cooperate for conversions (or some other coordination in userspace). If >> private memory gets converted to shared, iommufd needs to know to remap >> the pages as shared in the IO page tables. >> >> [4] https://lore.kernel.org/all/[email protected]/ >> >> > • Revocation (which I've tested correctly breaks down large pages in >> > both EPT/NPT and IOMMU). >> >> See above. >> I missed considering breaking down huge pages (the folios). guest_memfd will definitely need to cooperate with the provider to break down the pages. I'll need to prototype support with HugeTLBfs, probably after LPC. There I'll need to split pages on conversion to shared and merge on conversion to private. I think you meant breaking down the mappings for the large pages in stage 2? guest_memfd tells KVM to do that in one of the earlier prototypes for guest_memfd HugeTLB pages. >> > • AsyncPF support for pages requested by KVM. >> >> Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness. >> KVM MMU checks that the page is not there, then since async was >> configured, KVM tells the guest to schedule in something else. >> >> guest_memfd doesn't have support for async #PFs now. I imagine it would >> be something like KVM MMU firing off a kvm_gmem_get_pfn() request >> asynchronously, and then guest_memfd gets the page or PFN and inserts it >> in guest_memfd filemap, and then notifies KVM MMU? >> >> Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from >> the provider. >> >> I think it's orthogonal because guest_memfd could >> have gotten a PAGE_SIZE page (existing functionality) or gotten a page >> from the provider. > > Kind of, but I'm focusing on the guest_memfd provider interface, and > that needs to support the asychronous mode: asked for a PFN for a given > guest address, it returns -EAGAIN and then provides it later, and the > guest gets the right asyncpf behaviour: > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f I thought the same .alloc_folio would be used, except for an async #PF the entire kvm_gmem_get_folio(), which calls .alloc_folio() would be done in some background thread, so the provider doesn't have to have a different .alloc_folio_async(). If the provider does need to know, then the corresponding thing in this series could be: struct guest_memfd_provider_operations { void *(*attach)(struct file *resource_file); void (*release)(void *provider); struct folio *(*alloc_folio)(void *provider, pgoff_t index, - struct mempolicy *mpol); + struct mempolicy *mpol, bool nowait); void (*invalidate_folio)(void *provider, struct folio *folio); }; something like that? Either way I think it would have to be a later series after the first one (for upstreaming).

