One source, two GPU vendors, and memory bandwidth still gets the last word in this OpenMP LWE KEM paper.
trending12
01 02 Hybrid CPU-GPU sumcheck backend for Spartan provers, keeps polynomial state on GPU until it stops paying rent.03 Does AMD ROCm/HIP finally do Goldilocks STARK proving? The post includes qingming-stark-g64, CLI flow, and SCALE20-27 benchmarks.04 A zkML explainer on proving inference ran correctly, with circuits, quantization, and the usual proof-generation pain.05 LibFHE: A Numba-Based CUDA-Python Library for Non-RNS CKKS-BGV Fully Homomorphic Encryption on GPUs arxiv.orgPaper on LibFHE, a Numba/CUDA-Python GPU FHE library for non-RNS CKKS-BGV, because C++ was too mainstream.06 LibFHE: A Numba-Based CUDA-Python Library for Non-RNS CKKS-BGV Fully Homomorphic Encryption on GPUs arxiv.orgNon-RNS CKKS-BGV on GPUs, because apparently someone missed the usual RNS memo and still got competitive performance.07 LibFHE: A Numba-Based CUDA-Python Library for Non-RNS CKKS-BGV Fully Homomorphic Encryption on GPUs arxiv.orgNon-RNS CKKS-BGV on GPUs, apparently via Numba, and the paper claims C++-library-level performance without the C++.08 Native Goldilocks/G64 NTTs on an RTX4090, with qingming_fast/standard benchmarks and validated LDE domains up to 2^30.09 Qingming-g64-ntt: Native Goldilocks/G64 GPU NTT at 2^27 on RX 7900 XTX, and a reproducible benchmark plan ethresear.ch2^27 NTT on an RX 7900 XTX, plus a benchmark plan and correctness checks, for those who enjoy measuring crypto pain.10 End-to-end confidential AI for CPU+GPU TEEs, with attestation and benchmarks, not the usual single-box secure inference demo.11 Trimmed ICICLE repo for MSM, NTT, and ECNTT on BN254/BLS12-381, with CPU/CUDA backends and C++, Rust, Go bindings.12 Can you brick a CUDA collective by flipping its masks, lanes, or epochs? This paper says yes, then adds CICs.