Paper audits input-side jailbreak defenses for local LLMs, then shows semantic attacks happily walk past the formal guarantees.
trending2
01 02 Black-box attacks reconstruct secrets from LLM context windows, including SSNs, because refusal is apparently not a boundary.