Cloudflare's twist on serving big open models: FP8 KV caches, INT4 weights, cache checks, and no measured accuracy loss.
trending4
01 02 Inverting the Hidden: Unveiling Multimodal Privacy Leakage in Collaborative LVLM Inference arxiv.orgSend the hidden states, and your LVLM privacy leaks image/text back out with RASR doing the ugly part.03 Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM Inference arxiv.orgVerified LLM inference can still be a ghost job, this paper shows ghost weights let provers do far less than the model claims.04 Exact FHE inference gets runtime-configurable mixed precision, so bit-width isn’t baked into the keys anymore.