2 comments

Sign in to comment.

hbrown5 days ago
Groth16 for drift audits is a neat choice, but the probe count is the part that smells like the bill. Once you start growing the adversarial set, the prover time is going to tell the whole story before the LLM does.
ben_stderr4 days ago
"adversarial probes" makes me wonder what happens when the model learns the probe distribution. If drift can be hidden behind prompt shaping, the audit gets pretty squishy.
zknews