Hello Gregory, Thank you for the careful read of the documentation. I will rework the sections you flagged as below:
- Address model I will update the wording to mention that VMM does know the GPA: it builds the guest's CFMWS window and so chooses the guest-physical range the device can inhabit. What the host owns is the HPA and the placement within it, not knowledge of the GPA. I will add the two models you described, fixed placement for an accelerator that needs exact physical placement versus no fixed placement for pooled or compressed memory, with a small HPA/GPA example, and note that the current series under discussion implements the fixed-offset case. - Virtual decoders I will rename the The "Guest decoder and commit" section to "Virtual decoders" and add a short flow showing a guest vdecoder write being absorbed and a live read returning COMMITTED from the host-locked physical decoder. - "The kernel virtualizes ..." I will update it to say vfio-pci-core (its config-space permission hooks). - guest IOAS, stage-2 mapping, and struct-page-less coherent memory I will define these terms where first used and include that the range is handed to the device whole and never onlined as system RAM. I will also add details why a stale stage-2 mapping must not outlive the HDM window, and how the VMM rebuilds it. - Reset I will reword this section. FLRs are virtualized so a guest reset can't inflict physical CXL.mem/decoder effects on the host rather than loose wording "must not take an FLR" I can share the revised text ahead of the v6 posting if that is easier to review. Thanks, Manish > -----Original Message----- > From: Gregory Price <[email protected]> > Sent: Thursday, September 17, 2026 1:03 AM > To: Manish Honap <[email protected]> > Cc: [email protected]; [email protected]; Ankit Agrawal <[email protected]>; > [email protected]; [email protected]; [email protected]; > Srirangan Madhavan <[email protected]>; [email protected]; > [email protected]; [email protected]; [email protected]; > [email protected]; [email protected]; [email protected]; Yishai > Hadas <[email protected]>; Shameer Kolothum Thodi > <[email protected]>; [email protected]; [email protected]; > [email protected]; [email protected]; [email protected]; Neo Jia > <[email protected]>; Krishnakant Jaju <[email protected]>; Vikram Sethi > <[email protected]>; Zhi Wang <[email protected]>; linux- > [email protected]; [email protected]; [email protected]; > [email protected]; [email protected]; linux- > [email protected]; [email protected] > Subject: Re: [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type- > 2 device passthrough > > External email: Use caution opening links or attachments > > > On Thu, Sep 17, 2026 at 12:05:39AM +0530, [email protected] wrote: > > From: Manish Honap <[email protected]> > > > > 1) Thank you so much for writing documentation, i truly appreciate this. > > 2) I apologize in advance for my terseness, I know writing is hard, > please do not interpret this as disliking your writing or series. > > > +Address model > > +============= > > + > > +The HDM memory is a coherent host physical range (HPA). The host > > +kernel resolves that range before the guest sees the device, and owns > > +it for the bind lifetime. The guest only chooses where the memory > > +appears in its own physical address space (GPA), by programming a > > +virtual endpoint HDM decoder. The guest never reprograms the physical > decoder. > > + > > +The kernel holds the HPA and does not see the GPA. The guest programs > > +a GPA and does not see the HPA. The VMM holds the device fd, reads > > +the committed base from the decoder-register region described below, > > +and maps the HPA-backed HDM region at the GPA the guest committed. > > +The base the guest reads back is the GPA, not the HPA. > > + > > I think this must be slightly inaccurate / imprecise wording. > > The host *must* provide some form of physical memory window to the guest > at initialization time, otherwise the guest has no way to know - at boot time > - > that there's even a window of memory it can use. > > That's what the CFMWS is. This is initialized by the hypervisor - which is > controlled by the host. > > So the host (at least the VMM) must know, for the region the entire device > *could* inhabit, what that GPA is - because it's the one that makes the CFMWS > for the guest. > > If this is not the case, then something is missing from this documentation to > explain why. > > > If you're actually trying to say is that the GPA's programmed into the virtual > decoders are largely symbolic - this at best feels a bit inaccurate and > simply an > implementation detail. > > The host's virtio device could enforce ....: > > Host Range: > CFMWS HPA - [0x10000, 0x20000] > | | > Guest Range: | | > CFMWS GPA - [0x50000, 0x60000] > > In that case, you'd get the following translation... > vdecoder0.0 - [0x58000, 0x60000] > CFMWS GPA - [0x58000, 0x60000] > CFMWS HPA - [0x18000, 0x20000] > > Or the virtio device could not enforce that and let the host page-fault just > hand > it a random page from the actual CXL device. > > vdecoder0.0 - [0x58000, 0x60000] > CFMWS GPA - [0x58000, 0x60000] > | > No discrete host mapping > > > The former makes sense if the device (accelerator) requires exact physical > placement to do its accelerator nonsense. > > The latter makes sense if the device (accelerator) doesn't care about > placement > (compressed memory). > > This is not saying we need support both out of the box, but we shouldn't lock > ourselves into the former unless there's some reason why the latter is not > reasonable. > > Can you please help document what the actual expected behavior is with > examples in the Address model section so it's easier to understand the intent? > That will help quite a bit. > > > +Guest decoder and commit > > +======================== > > + > > +The guest programs its virtual endpoint decoder through the trapped > > +region: it writes a base (a GPA), a size, and then the COMMIT bit. > > +The host already resolved and committed the physical placement before > > +the guest ran, > > So the host does know GPA, just not exact placement. > > > so a live read of the decoder always shows COMMITTED and the > > +guest's commit poll completes. The physical decoder is never > > +rewritten; the guest's writes are absorbed. > > + > > Rather clunky, round-about way to say "The guest decoders are > virtualized". If possible, it would be nice to formalize this concept > into "Virtual Decoders" - since that's what this is. > > With that concept i think you can probably generate some nice diagrams that > show how the guest vdecoder's interact with the host drivers. > > > +The VMM observes the commit, reads the committed base, and maps the > > +HDM region at that GPA. > > + > > So the host does know the GPA. > > > +CXL Device DVSEC > > +================ > > + > > +The kernel virtualizes the CXL Device DVSEC body through the > > +config-space permission hooks. Reads and writes inside the DVSEC body > > +use a per-open shadow; a guest write stays in the shadow and does not > reach hardware. > > +Accesses outside the DVSEC body go to the device as usual. > > + > > "The kernel" - what part? vfio-pci ? the vmm ? > > > > +DMA and iommufd > > +=============== > > + > > +A Type-2 accelerator issues ATS-translated DMA to addresses inside > > +its own HDM window, so that range must be present in the guest IOAS > > +that backs the nested stage-2 translation. > > Type-2, ATS, DMA, HDM window, guest IOAS, stage-2 translation > > I think the only thing i don't know in this sentence is "guest IOAS" and it's > still > hurting my brain to read. > > Are all accelerators expect to have this particular interaction, or just > yours? > > > The HDM range is struct-page-less coherent > > +memory, which a userspace-VA ``IOMMU_IOAS_MAP`` cannot pin. > > + > > The hardest part about writing about virtualization is keeping a consistent > mental model from section to section. > > which userspace? guest? host? (i presume guest here) > > `struct-page-less coherent memory` > e.g. the host never hotplugs this, it hands the entire region > directly to the VFIO device, right? > > I think this would be nice to spell out somewhere. > > > +The HDM memory region is therefore exportable as a dma-buf: > > +``VFIO_DEVICE_FEATURE_DMA_BUF`` on that region returns an fd that > > +iommufd maps with ``IOMMU_IOAS_MAP_FILE``, mapping the physical > range > > +without a VA or a page pin. The dma-buf is revoked whenever the > > +mapping is torn down (reset, power transition, teardown), so a stale > > +stage-2 mapping cannot outlive the HDM window. > > + > > For the sake of readers, I think either a little bit more information on this > "stage-2 mapping" concept is needed to make sense of what's going on here > and why it mustn't outlive the HDM window. > > > +Reset > > +===== > > + > > +A CXL Type-2 function must not take a Function Level Reset: an FLR > > +resets the coherent CXL.mem state and the HDM decoder. The PCI core > > +reflects this by preferring the CXL reset over FLR, so a function > > +reset of a CXL device runs the CXL DVSEC reset sequence, which resets > > +the function and then restores the HDM decoder and the PCI config state. > > + > > I think what you're trying to say is that FLRs are never passed to the device > because it can cause physical device effects that defeat the purpose of the > virtualization, yes? > > So we virtualize FLRs... > > > +A guest requests a reset by writing Initiate CXL Reset in the DVSEC. > > +That write only stamps completion in the shadow. The real reset runs > > +at the vfio reset points (the reset ioctl and a virtualized FLR through > > config > space): > > As you describe here. > > So it's not that an accelerator "must not take an FLR" - it's that FLRs are > virtualized to prevent deleterious effects on the host / hardware. > > Am I misunderstanding this? > > > +the kernel zaps the HDM mapping and revokes the dma-buf, then runs > > +the CXL reset, which always clears the device memory, and restores > > +and re-samples the decoder afterwards. A CXL port masks Secondary Bus > > +Reset by default, so a ``VFIO_DEVICE_PCI_HOT_RESET`` does not reach > > +the endpoint and the HDM state is untouched. If the port has SBR > > +unmasked the reset can decommit the decoder without restoring it, so > > +the reset_done handler gates HDM access; a ``VFIO_DEVICE_RESET`` then > > +runs the CXL reset sequence and restores it. > > + > > +The decoder register region is served by live reads of the hardware > > +decoder with guest writes absorbed: the decoder is committed and > > +locked by the host, so a guest can neither decommit nor reprogram it, > > +and the kernel keeps no shadow of the decoder state. > > This is basically what I said at the beginning - it must either be that the > host > provides locked auto-decoders at boot, or it must provide proper > virtualization > so that the decoders settings are fully virtualized. > > Seems it's the former, and that makes sense. Please correct me if i'm > misunderstanding. > > > After a reset the kernel restores and > > +re-samples the firmware-committed decoder, so the geometry the guest > > +reads back is unchanged. A VMM that dropped its HDM mapping, for > > +example across a reset or a D3hot->D0 transition, must rescan the > > +decoder and rebuild its > > +stage-2 mapping before it resumes HDM access. > > + > > Yeah i think we need a bit more information about this stage-2 mapping > rebuild to make sense of this. Maybe I'm just not read-up enough on this > particular setup - is there another part of the docs you can link to that talk > about this, or are you able to share some details as to what this rebuild > process looks like? > > ~Gregory

