On Mon, Sep 21, 2026, 8:21 PM Niels Möller nisse@lysator.liu.se wrote:
Niels Möller nisse@lysator.liu.se writes:
- Probably remove the ml_kem_params from the api, and instead use separate functions for ML-KEM 768 and ML-KEM 1024.
I've added specific functions in the most simple way on the branch refactor-ml-kem, but not yet un-exported the ml_kem params things.
I've now merged the changes on the refactor-ml-kem branch. Deletes the generic api (and makes the ml_kem_params struct completely internal). I've also done a couple of optimization, with speedup of 30%-50% when benchmarking on x86_64.
I think the implementation is already fast enough plus the optimizations are not on the hot path so it doesn't contribute much to the overall performance. Also, the new changes haven't been run yet on CI which keeps the other MRs open until this one closed.
- Add functions exposing structs for expanded keys. To avoid having to expand them over and over again if using the same key repeatedly. And for all-in-one functions, these structs could also be used as the types for needed scratch space,
After looking a bit closer, I'm not sure this is a good idea. We'll see.
I ran the changes on different CI jobs and the pattern passes successfully each time for the entire pipeline. One needs to trigger the CI first to decide.
[ ... key generation ...] Another observation is that need for scratch space can be generally reduced by not computing and storing values long before they are needed. The A matrix is used for one matrix x vector multiplication, A * s. If s is generated upfront, it's sufficient to compute one row of A at a time or even a single scalar element at a time). Similarly, noice vector e could be generated and applied one element at a time.
[...]
Encapsulation uses a slightly different matrix x vector multiplication, A^T * y, so here it would be natural to generate A one column at a time.
On the other hand, top-level decapsulation uses both low-level decapsulation and encapsulation, so it needs A twice.
The last sentence seems wrong, low-level decapsulation (inner_decrypt) doesn't use the matrix A at all.
You are looking at this very closely, but there is a subtle structural detail in how ml-kem handles its keys that changes how often matrix A actually needs to be generated. in some optimized implementations, decaps doesn't use the matrix A but in this case it does need A twice without considering the initial run. The sentence is correct as it stands.
So for the three top-level functions generate, encap and decap, all of them use A only once, and each element of A only once. So it should be straight forward to generate them one at a time. In principle, it should work fine to generate only 2 (out of 256) uint16_t coeffients at a time, but I suspect that going that far would come with a measurable performance penalty.
the caller appears to ignore the "performance penalty" since it's no longer an issue on modern devices. Running those parameters on non constrained environment is no brainer for developers as long they are serious about having one clean API.
best regards, Mamoun
Regards, /Niels
-- Niels Möller. PGP key CB4962D070D77D7FCB8BA36271D8F1FF368C6677. Internet email is subject to wholesale government surveillance. _______________________________________________ nettle-bugs mailing list -- nettle-bugs@lists.lysator.liu.se To unsubscribe send an email to nettle-bugs-leave@lists.lysator.liu.se