On Tue, 2026-09-29 at 16:41 -0700, Ackerley Tng wrote: > David Woodhouse <[email protected]> writes: > > > On Mon, 2026-09-28 at 15:59 -0700, Ackerley Tng wrote: > > > David Woodhouse <[email protected]> writes: > > > > > > > On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote: > > > > > > > > > > === What if my provider doesn't deal with pages? > > > > > > > > > > Frank and David Woodhouse [2] have use cases for providers that use > > > > > PFNs, and this is definitely something a generic interface must > > > > > support. > > > > > > > > > > Here are some options I can think of: > > > > > > > > > > 1. Don't define .alloc_folio(), instead define .alloc_pfn() > > > > > 2. Refactor .alloc_folio() to .alloc_pfn() > > > > > > > > > > Either way, I think guest_memfd should still be the primary manager of > > > > > the memory. > > > > > > > > Thanks for working on this. > > > > > > > > > > Thanks for the quick reply! > > > > > > > I'm not sure what you mean by 'primary manager' here but my use case is > > > > > > Naming is hard! See comment on revocation below. > > > > > > > for the provider to be the ultimate arbiter of who owns which PFN, and > > > > revocation of the same. Think of it like a filesystem backed by > > > > external memory. Each guest's memory is a file, and a guest can > > > > *donate* specific pages of its memory to other guests (which in the > > > > > > Putting donation aside first - I don't have a good picture of how a > > > guest would tell the host that it's willing to share a page - can't > > > comment. Would love to find out more. > > > > > > > context of the *interface* only means that we need revocation). > > > > That's implementation-specific. Think of things like Nitro Enclaves > > where a guest 'donates' its memory back to the hypervisor to be used by > > another microvm. But that interface isn't the guest_memfd concern; as I > > said, in the context of the guest_memfd interface, there is only > > revocation: that page went away, whatever the reason. > > > > Okay we're on the same page then. :) guest_memfd only needs to know that > the page went away, guest_memfd cannot say no. > > > > Going with this revocation part, beginning with a more basic use case: > > > I'm thinking that the provider can notify guest_memfd that the page or > > > PFN is going away. > > > > Right. That part I have working in what I was posting. > > [5] > > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf06 > > > > That is similar to what I had in mind except the parameters, I was > thinking inode, offset, size. > > kvm_gmem_invalidate_range() in your patch [5] does have a kvm_gmem > prefix, but it's really a direct call to tell KVM MMU to zap the gfn > range, guest_memfd doesn't get to update its own state.
Late here now, but it occurs to me that I need to check if that makes sense. Does the provider *know* the GFN, or only the offset within its own object — which might *not* be in a memslot starting at GFN0? Tangentially: I'd actually quite like the provider to *know* about any such offset, because if possible I'd like it to be able to be opinionated about where in the guest pages are mapped. > > > > > > Yup, guest_memfd needs some way to track PFNs in addition to folios. I > > > think for tracking we can reference DAX, the part I haven't looked at is > > > how to let guest_memfd handle events. Like if there's memory failure on > > > a PFN, how do we let guest_memfd handle it? > > > > Telling the guest about correctable and uncorrectable errors, you mean? > > I would have thought that's up to the VMM, until the point where (see > > revocation)? > > > > Informing the guest is definitely the userspace VMM's action to > take. > > guest_memfd's role here is updating it's own tracking, and perhaps > splitting folios. When guest_memfd supports huge pages, one thing high > on the wishlist is to split the huge folio and zap only the page where > there was an error and not the huge page. > > Advantage of splitting: If the failure happened in the last 4K, the > first 1G - 4K can continue to be used by the guest. If the guest never > touches the last 4K, it doesn't have to know about the failure at all. Right, but this is still just revocation of a single 4KiB page and expecting both KVM and IOMMU to shatter the large page accordingly, isn't it? That much I already had working. (The IOMMU shatter does need to be atomic and miss-less, which I don't think is the case on Arm but could be.) > > > > I think where the memory comes *from* is an implementation detail for > > the provider. It provides a PFN (or folio, if you must). All else is > > not the business of KVM or IOMMUFD. > > > > So we already have > > dmabuf --wraps--> CMA (and others) > > We could have > > [a] guest_memfd --wraps--> dmabuf --wraps--> CMA > > or > > [b] guest_memfd --wraps--> CMA > > I agree guest_memfd shouldn't care where the PFN comes from, but this > does impact where the provider_ops code is implemented though. > > > If we go with [b], CMA needs to find a non-dmabuf way for userspace to > get some resource fd. If we go with [a], dmabuf folks need to support > this :) I'm not sure I see how CMA is relevant to the *interface*. A given *implementation* (provider) of a guest_memfd resource might use CMA for its backing store. Or might use hugetlbfs. Or something DAX- like. Once the implementation has provided its guest_memfd operations with its 'get_pfn' method, nobody else cares *where* it finds the PFNs that it returns when we invoke that method. > I missed considering breaking down huge pages (the folios). guest_memfd > will definitely need to cooperate with the provider to break down the > pages. KVM manages this part on its own. The guest_memfd implementation only has to tell it to revoke. It's been a while, but I don't even remember having to do anything special. https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=ee95ed5aaf0 > I'll need to prototype support with HugeTLBfs, probably after LPC. There > I'll need to split pages on conversion to shared and merge on conversion > to private. > > I think you meant breaking down the mappings for the large pages in > stage 2? guest_memfd tells KVM to do that in one of the earlier > prototypes for guest_memfd HugeTLB pages. > > > > > • AsyncPF support for pages requested by KVM. > > > > > > Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness. > > > KVM MMU checks that the page is not there, then since async was > > > configured, KVM tells the guest to schedule in something else. > > > > > > guest_memfd doesn't have support for async #PFs now. I imagine it would > > > be something like KVM MMU firing off a kvm_gmem_get_pfn() request > > > asynchronously, and then guest_memfd gets the page or PFN and inserts it > > > in guest_memfd filemap, and then notifies KVM MMU? > > > > > > Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from > > > the provider. > > > > > > I think it's orthogonal because guest_memfd could > > > have gotten a PAGE_SIZE page (existing functionality) or gotten a page > > > from the provider. > > > > Kind of, but I'm focusing on the guest_memfd provider interface, and > > that needs to support the asychronous mode: asked for a PFN for a given > > guest address, it returns -EAGAIN and then provides it later, and the > > guest gets the right asyncpf behaviour: > > https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34f > > I thought the same .alloc_folio would be used, except for an async #PF > the entire kvm_gmem_get_folio(), which calls .alloc_folio() would be > done in some background thread, so the provider doesn't have to have a > different .alloc_folio_async(). > > If the provider does need to know, then the corresponding thing in this > series could be: > > struct guest_memfd_provider_operations { > void *(*attach)(struct file *resource_file); > void (*release)(void *provider); > struct folio *(*alloc_folio)(void *provider, pgoff_t index, > - struct mempolicy *mpol); > + struct mempolicy *mpol, bool nowait); > void (*invalidate_folio)(void *provider, struct folio *folio); > }; > > something like that? Yeah, the 'nowait' argument on get_pfn() is how I have it working at the moment. And then, as you say, KVM calls it again *without* nowait, in a context which can sleep. > Either way I think it would have to be a later series after the first > one (for upstreaming). I don't have a particular use case for asyncpf right now personally, but I think we *should* get the design for the interface right. We should do it on the IOMMU side too, because ATS+PRI exists.
smime.p7s
Description: S/MIME cryptographic signature

