Polynomial-time model extraction now reaches multi-head softmax attention, with an 8D 6-head toy recovered to near machine precision.
trending2
01 02 Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM Inference arxiv.orgVerified LLM inference can still be a ghost job, this paper shows ghost weights let provers do far less than the model claims.