Native Goldilocks/G64 NTTs on an RTX4090, with qingming_fast/standard benchmarks and validated LDE domains up to 2^30.
trending7
01 02 Qingming-g64-ntt: Native Goldilocks/G64 GPU NTT at 2^27 on RX 7900 XTX, and a reproducible benchmark plan ethresear.ch2^27 NTT on an RX 7900 XTX, plus a benchmark plan and correctness checks, for those who enjoy measuring crypto pain.03 High-Performance NTT Accelerators for PQC leveraging Unified Redundant Arithmetic and Fine-Tuned Microarchitecture arxiv.orgFPGA NTT/INTT accelerators for ML-KEM/ML-DSA, with redundant arithmetic and tidy microarchitecture to shave cycles.04 Trimmed ICICLE repo for MSM, NTT, and ECNTT on BN254/BLS12-381, with CPU/CUDA backends and C++, Rust, Go bindings.05 Paper on tweaking TPU-style systolic arrays to run FHE NTTs with low-precision math and full-precision reconstruction.06 MPX folds direct polynomial multiply into a matrix systolic array, sidestepping NTT on the same silicon for crypto workloads.07 FEnc$^2$: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment Encoding arxiv.orgPaper on CKKS packing for private inference, using fragment encoding to cut rotations and ciphertexts. Packing, but make it baroque.