On 9/21/26 1:42 AM, Ancheng.Qiao wrote:
With Zcmp (cm.push) and -Os, the prologue folds the callee-saved FPRs
(fs0..fs11, i.e. f8/f9/f18..f31) into the stack region that cm.push
reserves using get_multi_push_fpr_mask (multi_push_additional /
UNITS_PER_WORD). That region is measured in bytes, but a double is
8 bytes (UNITS_PER_FP_REG), so dividing by the word size (4 bytes)
lets the mask fold up to twice as many FPRs as the reserved space can
actually hold.
The surplus FPRs (typically fs6..fs11, i.e. f22..f27) therefore end up
stored at negative offsets of the current SP, with the stack allocation
that should cover them deferred to a trailing "addi sp,sp,-N" that
comes after the stores. This is a use-the-stack-before-allocating-it
defect: callee-saved FP state lives in the not-yet-allocated region
below SP, which an asynchronous interrupt running on the interrupted
SP (e.g. a nested CLIC/Zcswl handler) can overwrite before the frame is
claimed, silently corrupting f22..f27.
gcc/ChangeLog:
* config/riscv/riscv.cc (riscv_expand_prologue): Fold into the
cm.push reserve only as many callee-saved FPRs as it can hold by
dividing the additional reserve by UNITS_PER_FP_REG.
(riscv_expand_epilogue): Likewise.
gcc/testsuite/ChangeLog:
* gcc.target/riscv/zcmp_fpr_save_below_sp.c: New test. Compile an
FP-register-heavy function at -Os with Zcmp and check that no
callee-saved FPR is saved at a negative offset of SP.
Signed-off-by: Ancheng.Qiao<[email protected]>
Thanks. I made the test only run on rv32 and pushed the patchkit to the
trunk.
Jeff