https://github.com/arsenm commented:
It seems this instruction wasn't exactly "new" in gfx94*. It used to have a
slightly different name, and had fewer control bits. This area seems to have a
lot of churn, this has happened a few times with the flushes. Would you be
better served by using an abstract fence instead? Here's the comparative
codegen:
```
┌────────────────────┬─────────────────────────────────────────────────────────┬───────────────────────────────────────────────────┐
│ Instruction │ Builtin
│ IR │
├────────────────────┼─────────────────────────────────────────────────────────┼───────────────────────────────────────────────────┤
│ buffer_inv sc0 │ __builtin_amdgcn_fence(__ATOMIC_ACQUIRE, "workgroup",
│ fence syncscope("workgroup") acquire, !mmra !0 in │
│ (work-group) │ "global") + tgsplit
│ a function with "amdgpu-tg-split" │
├────────────────────┼─────────────────────────────────────────────────────────┼───────────────────────────────────────────────────┤
│ buffer_inv sc1 │ __builtin_amdgcn_fence(__ATOMIC_ACQUIRE, "agent",
│ fence syncscope("agent") acquire, !mmra !0 │
│ (agent) │ "global")
│ │
├────────────────────┼─────────────────────────────────────────────────────────┼───────────────────────────────────────────────────┤
│ buffer_inv sc0 sc1 │ __builtin_amdgcn_fence(__ATOMIC_ACQUIRE, "", "global")
│ fence acquire, !mmra !0 │
│ (system) │
│ │
└────────────────────┴─────────────────────────────────────────────────────────┴───────────────────────────────────────────────────┘
```
https://github.com/llvm/llvm-project/pull/228497
_______________________________________________
cfe-commits mailing list
[email protected]
https://lists.llvm.org/cgi-bin/mailman/listinfo/cfe-commits