On Mon, Sep 21, 2026 at 8:21 PM Niels Möller nisse@lysator.liu.se wrote:
Niels Möller nisse@lysator.liu.se writes:
- Probably remove the ml_kem_params from the api, and instead use separate functions for ML-KEM 768 and ML-KEM 1024.
I've added specific functions in the most simple way on the branch refactor-ml-kem, but not yet un-exported the ml_kem params things.
I've now merged the changes on the refactor-ml-kem branch. Deletes the generic api (and makes the ml_kem_params struct completely internal). I've also done a couple of optimization, with speedup of 30%-50% when benchmarking on x86_64.
I'm trying to understand the optimization logic here to see if there is a balance we can achieve. What specific speedup do you intend to target?
best regards, Mamoun
- Add functions exposing structs for expanded keys. To avoid having to expand them over and over again if using the same key repeatedly. And for all-in-one functions, these structs could also be used as the types for needed scratch space,
After looking a bit closer, I'm not sure this is a good idea. We'll see.
[ ... key generation ...] Another observation is that need for scratch space can be generally reduced by not computing and storing values long before they are needed. The A matrix is used for one matrix x vector multiplication, A * s. If s is generated upfront, it's sufficient to compute one row of A at a time or even a single scalar element at a time). Similarly, noice vector e could be generated and applied one element at a time.
[...]
Encapsulation uses a slightly different matrix x vector multiplication, A^T * y, so here it would be natural to generate A one column at a time.
On the other hand, top-level decapsulation uses both low-level decapsulation and encapsulation, so it needs A twice.
The last sentence seems wrong, low-level decapsulation (inner_decrypt) doesn't use the matrix A at all.
So for the three top-level functions generate, encap and decap, all of them use A only once, and each element of A only once. So it should be straight forward to generate them one at a time. In principle, it should work fine to generate only 2 (out of 256) uint16_t coeffients at a time, but I suspect that going that far would come with a measurable performance penalty.
Regards, /Niels
-- Niels Möller. PGP key CB4962D070D77D7FCB8BA36271D8F1FF368C6677. Internet email is subject to wholesale government surveillance. _______________________________________________ nettle-bugs mailing list -- nettle-bugs@lists.lysator.liu.se To unsubscribe send an email to nettle-bugs-leave@lists.lysator.liu.se