Hi Shiwei, Ordering is guaranteed by the fence either way, asm or C makes no difference.
For the generated code, the compiler cannot see through inline asm, so the 8/16/32-bit reads end up with an extra extension instruction (tested with gcc 16 and clang 22). > lbu/lhu/lwu/ld for reads (zero-extending to avoid sign-bit pollution) The compiler picks the right one from the context, there is no pollution. Overall I don't see a real benefit here and would rather keep the generic implementation. What do you think?

