Hello Gregory,

Thank you for the careful read of the documentation. I will rework the
sections you flagged as below:

- Address model

  I will update the wording to mention that VMM does know the GPA: it builds
  the guest's CFMWS window and so chooses the guest-physical range the device
  can inhabit. What the host owns is the HPA and the placement within it, not
  knowledge of the GPA.
  I will add the two models you described, fixed placement for an accelerator
  that needs exact physical placement versus no fixed placement for pooled or
  compressed memory, with a small HPA/GPA example, and note that the current
  series under discussion implements the fixed-offset case.

- Virtual decoders

  I will rename the The "Guest decoder and commit" section to "Virtual
  decoders" and add a short flow showing a guest vdecoder write being absorbed
  and a live read returning COMMITTED from the host-locked physical decoder.

- "The kernel virtualizes ..."

  I will update it to say vfio-pci-core (its config-space permission hooks).

- guest IOAS, stage-2 mapping, and struct-page-less coherent memory

  I will define these terms where first used and include that the range is
  handed to the device whole and never onlined as system RAM. I will also add
  details why a stale stage-2 mapping must not outlive the HDM window, and how
  the VMM rebuilds it.

- Reset

  I will reword this section. FLRs are virtualized so a guest reset can't
  inflict physical CXL.mem/decoder effects on the host rather than loose
  wording "must not take an FLR"

I can share the revised text ahead of the v6 posting if that is easier to
review.

Thanks,
Manish

> -----Original Message-----
> From: Gregory Price <[email protected]>
> Sent: Thursday, September 17, 2026 1:03 AM
> To: Manish Honap <[email protected]>
> Cc: [email protected]; [email protected]; Ankit Agrawal <[email protected]>;
> [email protected]; [email protected]; [email protected];
> Srirangan Madhavan <[email protected]>; [email protected];
> [email protected]; [email protected]; [email protected];
> [email protected]; [email protected]; [email protected]; Yishai
> Hadas <[email protected]>; Shameer Kolothum Thodi
> <[email protected]>; [email protected]; [email protected];
> [email protected]; [email protected]; [email protected]; Neo Jia
> <[email protected]>; Krishnakant Jaju <[email protected]>; Vikram Sethi
> <[email protected]>; Zhi Wang <[email protected]>; linux-
> [email protected]; [email protected]; [email protected];
> [email protected]; [email protected]; linux-
> [email protected]; [email protected]
> Subject: Re: [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-
> 2 device passthrough
> 
> External email: Use caution opening links or attachments
> 
> 
> On Thu, Sep 17, 2026 at 12:05:39AM +0530, [email protected] wrote:
> > From: Manish Honap <[email protected]>
> >
> 
> 1) Thank you so much for writing documentation, i truly appreciate this.
> 
> 2) I apologize in advance for my terseness, I know writing is hard,
>    please do not interpret this as disliking your writing or series.
> 
> > +Address model
> > +=============
> > +
> > +The HDM memory is a coherent host physical range (HPA). The host
> > +kernel resolves that range before the guest sees the device, and owns
> > +it for the bind lifetime. The guest only chooses where the memory
> > +appears in its own physical address space (GPA), by programming a
> > +virtual endpoint HDM decoder. The guest never reprograms the physical
> decoder.
> > +
> > +The kernel holds the HPA and does not see the GPA. The guest programs
> > +a GPA and does not see the HPA. The VMM holds the device fd, reads
> > +the committed base from the decoder-register region described below,
> > +and maps the HPA-backed HDM region at the GPA the guest committed.
> > +The base the guest reads back is the GPA, not the HPA.
> > +
> 
> I think this must be slightly inaccurate / imprecise wording.
> 
> The host *must* provide some form of physical memory window to the guest
> at initialization time, otherwise the guest has no way to know - at boot time 
> -
> that there's even a window of memory it can use.
> 
> That's what the CFMWS is.  This is initialized by the hypervisor - which is
> controlled by the host.
> 
> So the host (at least the VMM) must know, for the region the entire device
> *could* inhabit, what that GPA is - because it's the one that makes the CFMWS
> for the guest.
> 
> If this is not the case, then something is missing from this documentation to
> explain why.
> 
> 
> If you're actually trying to say is that the GPA's programmed into the virtual
> decoders are largely symbolic - this at best feels a bit inaccurate and 
> simply an
> implementation detail.
> 
> The host's virtio device could enforce ....:
> 
> Host Range:
>    CFMWS HPA - [0x10000, 0x20000]
>                    |         |
> Guest Range:       |         |
>    CFMWS GPA - [0x50000, 0x60000]
> 
> In that case, you'd get the following translation...
>    vdecoder0.0 - [0x58000, 0x60000]
>    CFMWS GPA   - [0x58000, 0x60000]
>    CFMWS HPA   - [0x18000, 0x20000]
> 
> Or the virtio device could not enforce that and let the host page-fault just 
> hand
> it a random page from the actual CXL device.
> 
>    vdecoder0.0 - [0x58000, 0x60000]
>    CFMWS GPA   - [0x58000, 0x60000]
>                          |
>               No discrete host mapping
> 
> 
> The former makes sense if the device (accelerator) requires exact physical
> placement to do its accelerator nonsense.
> 
> The latter makes sense if the device (accelerator) doesn't care about 
> placement
> (compressed memory).
> 
> This is not saying we need support both out of the box, but we shouldn't lock
> ourselves into the former unless there's some reason why the latter is not
> reasonable.
> 
> Can you please help document what the actual expected behavior is with
> examples in the Address model section so it's easier to understand the intent?
> That will help quite a bit.
> 
> > +Guest decoder and commit
> > +========================
> > +
> > +The guest programs its virtual endpoint decoder through the trapped
> > +region: it writes a base (a GPA), a size, and then the COMMIT bit.
> > +The host already resolved and committed the physical placement before
> > +the guest ran,
> 
> So the host does know GPA, just not exact placement.
> 
> > so a live read of the decoder always shows COMMITTED and the
> > +guest's commit poll completes. The physical decoder is never
> > +rewritten; the guest's writes are absorbed.
> > +
> 
> Rather clunky, round-about way to say "The guest decoders are
> virtualized".   If possible, it would be nice to formalize this concept
> into "Virtual Decoders" - since that's what this is.
> 
> With that concept i think you can probably generate some nice diagrams that
> show how the guest vdecoder's interact with the host drivers.
> 
> > +The VMM observes the commit, reads the committed base, and maps the
> > +HDM region at that GPA.
> > +
> 
> So the host does know the GPA.
> 
> > +CXL Device DVSEC
> > +================
> > +
> > +The kernel virtualizes the CXL Device DVSEC body through the
> > +config-space permission hooks. Reads and writes inside the DVSEC body
> > +use a per-open shadow; a guest write stays in the shadow and does not
> reach hardware.
> > +Accesses outside the DVSEC body go to the device as usual.
> > +
> 
> "The kernel" - what part? vfio-pci ? the vmm ?
> 
> 
> > +DMA and iommufd
> > +===============
> > +
> > +A Type-2 accelerator issues ATS-translated DMA to addresses inside
> > +its own HDM window, so that range must be present in the guest IOAS
> > +that backs the nested stage-2 translation.
> 
> Type-2, ATS, DMA, HDM window, guest IOAS, stage-2 translation
> 
> I think the only thing i don't know in this sentence is "guest IOAS" and it's 
> still
> hurting my brain to read.
> 
> Are all accelerators expect to have this particular interaction, or just 
> yours?
> 
> > The HDM range is struct-page-less coherent
> > +memory, which a userspace-VA ``IOMMU_IOAS_MAP`` cannot pin.
> > +
> 
> The hardest part about writing about virtualization is keeping a consistent
> mental model from section to section.
> 
> which userspace? guest? host?  (i presume guest here)
> 
> `struct-page-less coherent memory`
>    e.g. the host never hotplugs this, it hands the entire region
>    directly to the VFIO device, right?
> 
>    I think this would be nice to spell out somewhere.
> 
> > +The HDM memory region is therefore exportable as a dma-buf:
> > +``VFIO_DEVICE_FEATURE_DMA_BUF`` on that region returns an fd that
> > +iommufd maps with ``IOMMU_IOAS_MAP_FILE``, mapping the physical
> range
> > +without a VA or a page pin. The dma-buf is revoked whenever the
> > +mapping is torn down (reset, power transition, teardown), so a stale
> > +stage-2 mapping cannot outlive the HDM window.
> > +
> 
> For the sake of readers, I think either a little bit more information on this
> "stage-2 mapping" concept is needed to make sense of what's going on here
> and why it mustn't outlive the HDM window.
> 
> > +Reset
> > +=====
> > +
> > +A CXL Type-2 function must not take a Function Level Reset: an FLR
> > +resets the coherent CXL.mem state and the HDM decoder. The PCI core
> > +reflects this by preferring the CXL reset over FLR, so a function
> > +reset of a CXL device runs the CXL DVSEC reset sequence, which resets
> > +the function and then restores the HDM decoder and the PCI config state.
> > +
> 
> I think what you're trying to say is that FLRs are never passed to the device
> because it can cause physical device effects that defeat the purpose of the
> virtualization, yes?
> 
> So we virtualize FLRs...
> 
> > +A guest requests a reset by writing Initiate CXL Reset in the DVSEC.
> > +That write only stamps completion in the shadow. The real reset runs
> > +at the vfio reset points (the reset ioctl and a virtualized FLR through 
> > config
> space):
> 
> As you describe here.
> 
> So it's not that an accelerator "must not take an FLR" - it's that FLRs are
> virtualized to prevent deleterious effects on the host / hardware.
> 
> Am I misunderstanding this?
> 
> > +the kernel zaps the HDM mapping and revokes the dma-buf, then runs
> > +the CXL reset, which always clears the device memory, and restores
> > +and re-samples the decoder afterwards. A CXL port masks Secondary Bus
> > +Reset by default, so a ``VFIO_DEVICE_PCI_HOT_RESET`` does not reach
> > +the endpoint and the HDM state is untouched. If the port has SBR
> > +unmasked the reset can decommit the decoder without restoring it, so
> > +the reset_done handler gates HDM access; a ``VFIO_DEVICE_RESET`` then
> > +runs the CXL reset sequence and restores it.
> > +
> > +The decoder register region is served by live reads of the hardware
> > +decoder with guest writes absorbed: the decoder is committed and
> > +locked by the host, so a guest can neither decommit nor reprogram it,
> > +and the kernel keeps no shadow of the decoder state.
> 
> This is basically what I said at the beginning - it must either be that the 
> host
> provides locked auto-decoders at boot, or it must provide proper 
> virtualization
> so that the decoders settings are fully virtualized.
> 
> Seems it's the former, and that makes sense.  Please correct me if i'm
> misunderstanding.
> 
> > After a reset the kernel restores and
> > +re-samples the firmware-committed decoder, so the geometry the guest
> > +reads back is unchanged. A VMM that dropped its HDM mapping, for
> > +example across a reset or a D3hot->D0 transition, must rescan the
> > +decoder and rebuild its
> > +stage-2 mapping before it resumes HDM access.
> > +
> 
> Yeah i think we need a bit more information about this stage-2 mapping
> rebuild to make sense of this.  Maybe I'm just not read-up enough on this
> particular setup - is there another part of the docs you can link to that talk
> about this, or are you able to share some details as to what this rebuild
> process looks like?
> 
> ~Gregory

Reply via email to