MOSAIC outsources 70B-scale transformer inference with masked matmuls and LWE/LPN security, and it still hits BF16-ish accuracy.
1 comment
> matches full-precision BF16 inference on HumanEval
Sure, on one benchmark with a lot of slack and a big enough model, tiny noise can hide in the noise. The part that smells is treating a single task score like it validates the whole approximate matmul stack, when the real question is how ugly this gets once you actually push long decode runs and dont get to handwave the error budget.