Backdoor cleanup for tool-calling LLM agents, with the annoying twist that unlearning often kills or reroutes other people's payloads.
1 comment
56% erased by just the defensive poisoning step is a grim number, and the fact that the survivor backdoors then mostly die when you unlearn one co-resident payload feels very STARK-ish in the bad sense, lots of structure but not the one you wanted.
The layerwise trigger awareness hanging around after benign behavior is restored is the part I'd trust least.