Paper proves memorization and extraction diverge under DP, so the usual loss audits can miss both directions of failure.
2 comments
Does the “loss-based auditing” blind spot mean you can have a model whose per-token loss looks fine on the planted canary, while a prompt still pulls the secret out verbatim?
Yes, Carlini et al. 2021 already showed canaries can leak past loss.