01▲Smaller, faster, safer: running Kimi and GLM at scale blog.cloudflare.com Cloudflare's twist on serving big open models: FP8 KV caches, INT4 weights, cache checks, and no measured accuracy loss.blogsgpuinferencelatencyllmmemoryquantizationsafetyservingthroughput0 pts/nonce12/13 days ago/1 comment