On 9/26/2026 7:33 PM, Akihiko Odaki wrote:
On 2026/09/24 18:55, Honglei Huang wrote:
Add userptr blob objects so the guest kernel can pin an existing
userspace mapping and advertise it as CREATE_BLOB backing entries.

- New virtio_gpu_object_userptr type for userptr resources
- Pin pages with pin_user_pages_fast() and FOLL_LONGTERM
- Omit FOLL_WRITE when VIRTGPU_BLOB_FLAG_USE_READONLY is set
- Charge FOLL_LONGTERM pins against RLIMIT_MEMLOCK
- DMA-map the scatterlist only when virtio_gpu_use_dma_api() is
   required; use DMA_TO_DEVICE for USE_READONLY blobs
- Mark writable pages dirty when unpinning
- Keep pages pinned until RESOURCE_UNREF is queued; drop them from
   cleanup_object() on the unref response or on create failure
- Clear userptr->pages on pin failure to avoid double-free on cleanup
- Reject unaligned or overflowing userptr ranges at create time
- Disallow PRIME export of userptr objects
- Save CREATE_BLOB params and restore userptr resources after
   hibernation without using the shmem restore path

The hibernation restore path maps with the same attrs and direction
used at create time, as documented next to the call. Like shmem
blobs, the DMA API can still bounce through a buffer there. This
design only avoids a second guest side shmem allocation and memcpy,
nothing more.

I don't see the restore path is documented though this says "documented next to the call".

Will fix in next version



...
+        goto err_cleanup;
+    }
+
+    userptr->sgt = sgt;
+
+    if (virtio_gpu_use_dma_api(vgdev->vdev)) {
+        enum dma_data_direction dir =
+            (userptr->flags & VIRTGPU_BLOB_FLAG_USE_READONLY) ?
+            DMA_TO_DEVICE : DMA_BIDIRECTIONAL;
+
+        ret = dma_map_sgtable(drm_dev_dma_dev(vgdev->ddev), sgt,
+                      dir, 0);

I checked the ROCm code and documentation, and I do not see how this mapping satisfies the HIP coherence contract when bounce buffers or explicit DMA cache maintenance are required.

HIP distinguishes two coherence models:

 > Coarse-grained coherence: The memory is considered up-to-date only
 > after synchronization performed using hipDeviceSynchronize(),
 > hipStreamSynchronize(), or any blocking operation that acts on the
 > null stream such as hipMemcpy(). To avoid the cache from being
 > accessed by a part of the system while simultaneously being written by
 > another, the memory is made visible only after the caches have been
 > flushed.
 >
 > Fine-grained coherence: The memory is coherent even while being
 > modified by a part of the system. Fine-grained coherence ensures that
 > up-to-date data is visible to others regardless of kernel boundaries.
 > This can be useful if both host and device operate on the same data.

https://rocm.docs.amd.com/projects/HIP/en/docs-10.0.0/how-to/ hip_runtime_api/memory_management/coherence_control.html

Fine-grained access requires coherence throughout execution; copying
between the original pages and a bounce buffer only at kernel boundaries
would not suffice. Coarse-grained access could support such copies, but
they must occur at the required synchronization points. I found no code
that synchronizes this DMA mapping at those points.


Whether dma_map_sgtable() bounces is decided by the platform's IOMMU/SWIOTLB behind the device, not by we request. attrs=0 here just matches the existing shmem blob path, and is exactly what native (non-virtualized) ROCm's own userptr support does too:

// drivers/gpu/drm/amd/amdgpu/amdgpu_amdkfd_gpuvm.c, kfd_mem_dmamap_userptr()
ret = dma_map_sgtable(adev->dev, ttm->sg, direction, 0);

And HIP's coarse-grained/fine-grained coherence isn't a DMA layer property, it is a GPU VM/cache thing. In upstream KFD it's implemented as GPU MMU page table attributes (MTYPE_CC/MTYPE_RW, the SNOOPED bit), set by whichever driver programs the real GPU's page tables for this memory:

// drivers/gpu/drm/amd/amdkfd/kfd_svm.c, svm_range_get_pte_flags()
mapping_flags |= coherent ? AMDGPU_VM_MTYPE_CC : AMDGPU_VM_MTYPE_RW;
pte_flags |= snoop ? AMDGPU_PTE_SNOOPED : 0;

I really want to disable the bounce buffers in DMA, but it is a platform behaviour, can not disable it in this layer. And in XEN, the dmapping for it is another form of iovector/simple sgtable. And if use the virtio-iommu or somthing else, the efficiency and the complexity of driver integration are both lower than using iovector/sgtable directly. And for how to handle the DMA thing for userptr, since we have so many concern about it, maybe we can drop this part in next version. But we may have the AI review warn/error.

Regards,
Honglei

Regards,
Akihiko Odaki

+        if (ret)
+            goto err_cleanup;
+
+        userptr->dma_dir = dir;
+        userptr->dma_mapped = true;
+    }
+
+    ret = virtio_gpu_userptr_get_entries(vgdev, userptr, &ents, &nents);
+    if (ret)
+        goto err_cleanup;
+
+    virtio_gpu_cmd_resource_create_blob(vgdev, &userptr->base, params, ents,
+                        nents);
+
+    userptr->base.params = *params;
+    virtio_gpu_add_object_to_restore_list(vgdev, &userptr->base);
+
+    *bo_ptr = &userptr->base;
+    return 0;
+
+err_cleanup:
+    virtio_gpu_cleanup_object(&userptr->base);
+    return ret;
+}



Reply via email to