SYNC_BO carries an offset and a size, but amdxdna_flush_bo() honours
them only on the vmap path: an imported BO is tested for first and
flushes its whole scatterlist, and the page-array fallback flushes every
page of the BO. A sync costs what the BO is worth rather than what the
caller asked to maintain.

amdxdna_gem_obj_vmap() maps an imported BO through dma_buf_vmap(), so
try the vmap path first and leave drm_clflush_sg() as the fallback for
an exporter that cannot serve one. Index the page-array fallback from
the requested offset. An imported BO now holds a kernel mapping from its
first sync until it is freed, as a shmem BO already does.

Measured on npu4, 64 MiB BO, pinned, minimum of 50 runs: an imported BO
cost 1056 us to sync at every size from 4 KiB up, and now tracks the
driver-owned BO at 0.7 us for 4 KiB, 17 us for 1 MiB and 1056 us for the
whole BO. The driver-owned column does not move.

An earlier version walked the scatterlist a page at a time instead. It
fixed the range case but cost about 179 ns per page of barrier and call
overhead, taking the whole-BO sync from 1056 to 3989 us, so this one
reuses the mapping instead.

This does not bracket the flush with dma_buf_begin_cpu_access() and
dma_buf_end_cpu_access(). The driver has never called them, here or
anywhere else, so the omission predates this change; what changes is that
the vmap path now serves an imported BO by default, which is where the
exporter's own coherency hook would matter most. Adding the bracket is
follow-up work rather than part of this one: the calls carry a direction
but no range, so pairing them with a ranged flush wants its own reasoning
and its own measurement.

Signed-off-by: Taimuraz Kaitmazov <[email protected]>
---
 drivers/accel/amdxdna/amdxdna_gem.c | 14 ++++++++++----
 1 file changed, 10 insertions(+), 4 deletions(-)

diff --git a/drivers/accel/amdxdna/amdxdna_gem.c 
b/drivers/accel/amdxdna/amdxdna_gem.c
index d0c846b02..4886f7c08 100644
--- a/drivers/accel/amdxdna/amdxdna_gem.c
+++ b/drivers/accel/amdxdna/amdxdna_gem.c
@@ -1234,6 +1234,8 @@ int amdxdna_drm_get_bo_info_ioctl(struct drm_device *dev, 
void *data, struct drm
 
 static int amdxdna_flush_bo(struct amdxdna_gem_obj *abo, u64 offset, u64 size)
 {
+       unsigned long first, nr_pages;
+       void *kva;
        u64 end;
 
        if (offset >= abo->mem.size)
@@ -1243,12 +1245,16 @@ static int amdxdna_flush_bo(struct amdxdna_gem_obj 
*abo, u64 offset, u64 size)
                return -EINVAL;
 
        size = min(abo->mem.size, end) - offset;
-       if (is_import_bo(abo))
+       first = offset >> PAGE_SHIFT;
+       nr_pages = (PAGE_ALIGN(offset + size) >> PAGE_SHIFT) - first;
+
+       kva = __amdxdna_gem_vmap(abo);
+       if (!IS_ERR(kva))
+               drm_clflush_virt_range(kva + offset, size);
+       else if (is_import_bo(abo))
                drm_clflush_sg(abo->base.sgt);
-       else if (amdxdna_gem_vmap(abo))
-               drm_clflush_virt_range(amdxdna_gem_vmap(abo) + offset, size);
        else if (abo->base.pages)
-               drm_clflush_pages(abo->base.pages, abo->mem.size >> PAGE_SHIFT);
+               drm_clflush_pages(&abo->base.pages[first], nr_pages);
        else
                return -EINVAL;
 
-- 
2.55.0

Reply via email to