On Wed, Sep 23, 2026 at 6:13 PM fengchengwen <[email protected]> wrote:
>
> >
> Hi Zhiping,
>
> On 8/1/2026 5:15 AM, Zhiping Zhang wrote:
> > Peer-to-peer DMA between a mlx5 NIC and a foreign PCIe endpoint
> > (typically a GPU or a vfio-pci passthrough device) traverses the host
> > PCIe fabric. The endpoint exporting the dma-buf knows which PCIe TLP
> > Processing Hint (TPH) Steering Tag yields the best placement for the
> > traffic it will sink: per-endpoint hint selection lets the root complex
> > or switch direct DMA to a specific cache slice / NUMA node, cutting
> > cross-socket snoop traffic and DRAM pressure under sustained p2p
> > workloads.
> >
> > Until now the mlx5 importer had no way to learn the exporter's chosen
> > ST tag, so dma-buf MRs were registered without TPH and ran with the
> > default (no-hint) routing. With dma_buf_get_pci_tph() in place this
> > patch wires up mlx5_ib to query that metadata at MR registration time
> > for p2p access and use it to program requester-side TPH on the outbound
> > mkey. If the exporter has no metadata, fall back to the existing
> > no-TPH path so behavior for non-TPH-aware exporters is unchanged.
>
>
> While working on v21 of the VFIO PCIe TPH series, we got review
> feedback that importers must add pci_p2pdma_distance() validation
> before retrieving ST values from a dma-buf.
>
> Your mlx5 patch follows the same importer pattern – it consumes
> dma-buf TPH metadata. I think this patch should also include the
> same distance check.
>
> Thanks,
> Chengwen
>

Hi Chengwen,

I don't think this belongs in the importer. Every in-tree dma-buf
caller of pci_p2pdma_distance() is the exporter, in its .attach,
clearing attach->peer2peer -- amdgpu_dma_buf.c, xe_dma_buf.c,
habanalabs memory.c. There is no importer-side caller.

mlx5 also can't be one: pci_p2pdma_distance() takes the exporter's
struct pci_dev, and an importer only has a struct dma_buf. Reaching it
would mean introducing PCI types into the dma-buf core, which is what
Christian's ack on patch 3 rules out.

The routing check is already enforced by the core before any p2p DMA
can happen: vfio_pci_dma_buf_map() goes through
dma_buf_phys_vec_to_sgt(), which switches on
pci_p2pdma_map_type(provider, attach->dev) and fails -EINVAL on
PCI_P2PDMA_MAP_NOT_SUPPORTED -- the same calc_map_type_and_dist()
computation pci_p2pdma_distance() performs. An ST that mlx5 programmed
into the mkey cannot reach the wire unless that mapping has succeeded.

Thanks,
Zhiping

Reply via email to