Recent work on embedded post-quantum cryptography has focused heavily on instruction-level optimization: improving arithmetic kernels, hand-tuning assembly, managing register allocation, and scheduling instructions. A new arXiv preprint argues that this focus misses a larger opportunity. Using ML-KEM on an Arm Cortex-M7 as a case study, the authors show that system-level factors—memory layout, function-call overhead, and compiler settings—can have a greater impact on performance than further tweaking the cryptographic kernels themselves.
The paper does not dismiss kernel-level work; rather, it situates it within a broader optimization space. By treating the full software stack as the target, the authors identify bottlenecks that instruction-level techniques alone cannot address. Their results indicate that careful attention to how the algorithm is integrated into the system—not just how its inner loops are written—can yield substantial speedups.
For developers of embedded PQC libraries, the takeaway is practical: before investing more effort in assembly micro-optimizations, consider the memory map, the calling convention, and the compiler's code generation. The study is a single case on a specific microcontroller, so the exact gains may not transfer directly to other cores or algorithms, but the direction is clear. System-level design deserves a place alongside kernel craftsmanship in embedded cryptography research.