On Thu, Sep 24, 2026 at 07:58:40AM +0000, Leonid Ravich wrote: > > 1. "Could you send me the patch so I can take a look?" (per-unit cost) > --------------------------------------------------------------------- > > Patch 4 is that patch -- the API-layer split, self-contained in > crypto/skcipher.c. The cost I measured, and where I believe it goes: > > An in-kernel microbench of skcipher_crypt_unit() against the legacy > per-unit loop (identical inner AES, VAES-AVX512, r7i.metal) shows a > fixed ~48-52 ns per data unit, CV <1%, and it is the *same* ~50 ns > for a 512 B unit and a 4096 B unit -- so it is per-call setup, not a > crypto effect. Against ~70 ns of VAES work for a 512 B sector that > is ~+70% on the crypto call itself; end-to-end in dm-crypt it is > ~2% of an ~18 us I/O and does not show up in fio at all (see the > Performance section below).
I think the issue is that we're generating the IV twice. Once in the Crypto API and once again in the DM layer. Not only is this slow, but it is actually wrong for decryption. You're ignoring the IVs on the disk. We really should only do it once, either in the DM layer or in the Crypto API. It should be transmitted via memory to the other side. My suggestion is to allocate memory for the IVs. Of course memory allocation can fail, but we have an easy fallback, which is to use the existing single-unit path. IOW if you succeed in allocating memory for storing the IVs, then invoke the multi-unit code path, otherwise fall back to the single-unit code path which iterates over the sectors one- by-one. To pass the IVs to the Crypto API (or back), just use the existing IV pointer and extend it by the number of units. Thanks, -- Email: Herbert Xu <[email protected]> Home Page: http://gondor.apana.org.au/~herbert/ PGP Key: http://gondor.apana.org.au/~herbert/pubkey.txt

