On Wed, Jul 15, 2026 at 06:47:26PM +0100, Matt Evans wrote:
> Add vfio_pci_dma_buf_find_pfn(), which a VMA fault handler can use to
> find a PFN.
> 
> This supports multi-range DMABUFs, which typically would be used to
> represent scattered spans but might even represent overlapping or
> aliasing spans of PFNs.
> 
> Because this is intended to be used in vfio_pci_core.c, we also need
> to expose the struct vfio_pci_dma_buf in the vfio_pci_priv.h header.
> 
> Signed-off-by: Matt Evans <[email protected]>
> ---
>  drivers/vfio/pci/vfio_pci_dmabuf.c | 153 ++++++++++++++++++++++++++---
>  drivers/vfio/pci/vfio_pci_priv.h   |  20 ++++
>  2 files changed, 160 insertions(+), 13 deletions(-)
> 
> diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c 
> b/drivers/vfio/pci/vfio_pci_dmabuf.c
> index c16f460c01d6..7c047400dfd1 100644
> --- a/drivers/vfio/pci/vfio_pci_dmabuf.c
> +++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
> @@ -9,19 +9,6 @@
>  
>  MODULE_IMPORT_NS("DMA_BUF");
>  
> -struct vfio_pci_dma_buf {
> -     struct dma_buf *dmabuf;
> -     struct vfio_pci_core_device *vdev;
> -     struct list_head dmabufs_elm;
> -     size_t size;
> -     struct phys_vec *phys_vec;
> -     struct p2pdma_provider *provider;
> -     u32 nr_ranges;
> -     struct kref kref;
> -     struct completion comp;
> -     u8 revoked : 1;
> -};
> -
>  static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
>                                  struct dma_buf_attachment *attachment)
>  {
> @@ -106,6 +93,146 @@ static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
>       .release = vfio_pci_dma_buf_release,
>  };
>  
> +int vfio_pci_dma_buf_find_pfn(struct vfio_pci_dma_buf *priv,
> +                           struct vm_area_struct *vma,
> +                           unsigned long fault_addr,
> +                           unsigned int order,
> +                           unsigned long *out_pfn)
> +{
> +     /*
> +      * Given a VMA (start, end, pgoffs) and a fault address,
> +      * search the corresponding DMABUF's phys_vec[] to find the
> +      * range representing the address's offset into the VMA, and
> +      * its PFN.
> +      *
> +      * The phys_vec[] ranges represent contiguous spans of VAs
> +      * upwards from the buffer offset 0; the actual PFNs might be
> +      * in any order, overlap/alias, etc.  Calculate an offset of
> +      * the desired page given VMA start/pgoff and address, then
> +      * search upwards from 0 to find which span contains it.
> +      *
> +      * On success, a valid PFN for a page sized by 'order' is
> +      * returned into out_pfn.
> +      *
> +      * Failure occurs if:
> +      * - A hugepage would cross the edge of the VMA,
> +      * - A hugepage isn't entirely contained within a range
> +      *   (including where it straddles the boundary between
> +      *   ranges),
> +      * - We find a range, but the final PFN isn't aligned to the
> +      *   requested order.
> +      *
> +      * Upon failure, -EAGAIN is returned and the caller is
> +      * expected to try again with a smaller order, which will
> +      * eventually succeed (order=0 will always work).
> +      *
> +      * It's suboptimal if DMABUFs are created with neighbouring
> +      * ranges that are physically contiguous, since hugepages
> +      * can't straddle range boundaries.  (The construction of the
> +      * ranges should merge them in this case.)
> +      *
> +      * Finally, vma_pgoff_adjust is used with a DMABUF created for
> +      * a VFIO BAR mmap: a BAR mapped with vm_pgoff > 0 creates a
> +      * DMABUF such that byte 0 of the VMA corresponds to byte 0 of
> +      * the DMABUF and byte 'vm_pgoff << PAGE_SHIFT' into the BAR.
> +      * To avoid double-offsetting in this scenario, subtracting
> +      * vma_pgoff_adjust from this (non-zero) vm_pgoff generates
> +      * the effective offset.
> +      */
> +
> +     const unsigned long pagesize = PAGE_SIZE << order;
> +     unsigned long vma_off = ((vma->vm_pgoff - priv->vma_pgoff_adjust) <<
> +                              PAGE_SHIFT) & VFIO_PCI_OFFSET_MASK;

Maybe I'm getting ahead of myself here.. but it seems like this
restricts us to only mapping DMABUFs at offsets < 1TB due to the 
VFIO_PCI_OFFSET_MASK (since we have HBMs on PCI devices now, hitting 1TB
may not be a very distant future).

While I understand this mask is needed to drop the BAR encoding in the
high bits. 

My worry is, if in the future a user were to export a massive
contiguous DMABUF (e.g., >1TB of aggregated HBM) and tried to mmap deep
into it (passing an offset >= 1TB), this bitwise AND would silently drop
the high bits, leading to silent data corruption.

I think we should explicitly reject such an mmap with -EINVAL like:

+const unsigned long pagesize = PAGE_SIZE << order;
+unsigned long vma_off = (vma->vm_pgoff - priv->vma_pgoff_adjust) << PAGE_SHIFT;

+/*
+ * Prevent silent wrap-around if the user mmaps a DMABUF at an
+ * offset greater than the VFIO index mask allows.
+ */
+if (unlikely(vma_off > VFIO_PCI_OFFSET_MASK))
+    return -EINVAL;

+vma_off &= VFIO_PCI_OFFSET_MASK;

> +     unsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);
> +     unsigned long rounded_page_end = rounded_page_addr + pagesize;
> +     unsigned long fault_offset;
> +     unsigned long fault_offset_end;
> +     unsigned long range_start_offset = 0;
> +     unsigned int i;
> +     int ret;
> +


Thanks,
Praan

Reply via email to