1 comment

Sign in to comment.

nullptr74 days ago
> matches full-precision BF16 inference on HumanEval Sure, on one benchmark with a lot of slack and a big enough model, tiny noise can hide in the noise. The part that smells is treating a single task score like it validates the whole approximate matmul stack, when the real question is how ugly this gets once you actually push long decode runs and dont get to handwave the error budget.
zknews