David Woodhouse <[email protected]> writes:

> On Fri, 2026-09-25 at 17:50 -0700, Ackerley Tng via B4 Relay wrote:
>>
>> === What if my provider doesn't deal with pages?
>>
>> Frank and David Woodhouse [2] have use cases for providers that use
>> PFNs, and this is definitely something a generic interface must
>> support.
>>
>> Here are some options I can think of:
>>
>> 1. Don't define .alloc_folio(), instead define .alloc_pfn()
>> 2. Refactor .alloc_folio() to .alloc_pfn()
>>
>> Either way, I think guest_memfd should still be the primary manager of
>> the memory.
>
> Thanks for working on this.
>

Thanks for the quick reply!

> I'm not sure what you mean by 'primary manager' here but my use case is

Naming is hard! See comment on revocation below.

> for the provider to be the ultimate arbiter of who owns which PFN, and
> revocation of the same. Think of it like a filesystem backed by
> external memory. Each guest's memory is a file, and a guest can
> *donate* specific pages of its memory to other guests (which in the

Putting donation aside first - I don't have a good picture of how a
guest would tell the host that it's willing to share a page - can't
comment. Would love to find out more.

> context of the *interface* only means that we need revocation).
>

Going with this revocation part, beginning with a more basic use case:
I'm thinking that the provider can notify guest_memfd that the page or
PFN is going away.

guest_memfd cannot say no, but it does have to tell KVM to zap the page
from stage 2, and for CoCo it may need to do more stuff.

The notification needs some pointer, so I was thinking that on .attach()
the provider would also note down which guest_memfd instance the
provider should notify.

The notifier could tell guest_memfd that some (offset, size) is going
away, and breaking down mappings in the EPT/NPT should be something
guest_memfd tells KVM to do, I think. Are you thinking that the provider
would directly break the mappings down in the EPT/NPT?

Does this notification mechanism work for you? IIUC regular DAX would
need it too, like if the device is unplugged for example.

> Should I be trying to rework parts of my series to live on top of this?

I hope to gather more feedback at LPC. It'd be nice to get early
feedback if you think it can work for you! The overall design still
needs feedback though, don't take this as the final direction.

> Things I have working but which I don't see here include:
>

This sounds like 4 or more different features! I think they do work with
this resource fd proposal.

>  • PFN support (which you mentioned).

Yup, guest_memfd needs some way to track PFNs in addition to folios. I
think for tracking we can reference DAX, the part I haven't looked at is
how to let guest_memfd handle events. Like if there's memory failure on
a PFN, how do we let guest_memfd handle it?

guest_memfd needs to at least be notified, so that it can handle the KVM
stuff: zapping and special stuff for CoCo. Whether the provider or
guest_memfd gets to handle it first can be worked out :)

>  • IOMMUFD support via dma-buf.

Is this about having guest_memfd work with dmabuf fds?

Fuad also brought up that Android uses dmabuf as an interface for CMA
memory. [3] IIUC the dmabuf fd is mmap()ed, then the userspace address
is set up in a memslot for the guest. To make this CMA memory
CoCo-friendly, we want to use it with guest_memfd.

Jason response was that guest_memfd should probably just directly get
memory from CMA [3].

I'm not super familiar with CMA or dmabufs, would either have a suitable
"resource fd" to pass to guest_memfd like how HugeTLBfs has a mount fd?
The "resource fd" should ideally have no way to mmap or read/write the
memory directly.

[3] 
https://lore.kernel.org/all/CA+EHjTxZ0N3Tfnid404B4tkb_E+Z8mODTHTgBPiF6=bwzp7...@mail.gmail.com/

Or is this about letting IOMMUFD get pages to map in IOMMU page tables
via an fd+offset instead of via mmap()-ed userspace addresses as in
IOMMU_IOAS_MAP_FILE, and that guest_memfd isn't a supported kind of fd
(yet)?

I think guest_memfd should support IOMMU_IOAS_MAP_FILE. One use case is
Confidential IO [4], where SNP would need to have private memory
(guest_memfd) mapped in the IO page tables.

The missing parts here are that guest_memfd and iommufd need to
cooperate for conversions (or some other coordination in userspace). If
private memory gets converted to shared, iommufd needs to know to remap
the pages as shared in the IO page tables.

[4] https://lore.kernel.org/all/[email protected]/

>  • Revocation (which I've tested correctly breaks down large pages in
>    both EPT/NPT and IOMMU).

See above.

>  • AsyncPF support for pages requested by KVM.

Is this kind of orthogonal? IIUC today KVM MMU handles the async-ness.
KVM MMU checks that the page is not there, then since async was
configured, KVM tells the guest to schedule in something else.

guest_memfd doesn't have support for async #PFs now. I imagine it would
be something like KVM MMU firing off a kvm_gmem_get_pfn() request
asynchronously, and then guest_memfd gets the page or PFN and inserts it
in guest_memfd filemap, and then notifies KVM MMU?

Within kvm_gmem_get_pfn(), guest_memfd should get the page or PFN from
the provider.

I think it's orthogonal because guest_memfd could
have gotten a PAGE_SIZE page (existing functionality) or gotten a page
from the provider.

Reply via email to