Anthropic's Claude Mythos Preview finds cracks in HAWK and 7-round AES-128, via a new cryptanalysis benchmark.
3 comments
anthropic's benchmark may be new, but the underlying game is the old biham and shamir differential cryptanalysis line, just with a model picking the branches instead of a human or a sat/SMt loop.
7-round aes-128 is still in the toy-attack zone, so i'm more curious whether this actually beats the usual automatic trail search from the mids 2000s or just packages it differently.
The 7-round AES bit is the only part here thats remotely easy to sanity check, and youre right to keep it in toy-attack land. But the leap from “model picked the branches” to “we found a cryptanalytic advance” is exactly the kind of marketing gloss that gets people clapping for a search heuristic with a bigger bill.
If Claude is just steering an old differential trail search, then the benchmark is measuring promptable brute force with nicer prose, not some new cryptanalytic capacity. I care a lot less that a model can rediscover a reduced-round attack than whether it does anything the mids 2000s tooling couldnt already do once you feed it enough compute and patience. HAWK sounds more interesting, but the post keeps waving at “expert level” and that smells like them trying to inflate one narrow success into a general claim
finding a crack in 7-round aes-128 and calling it a cryptanalysis benchmark feels like grading yourself on the warm-up set. hawk is the same story, a toy target with a fancy label, and somehow that still gets dressed up as progress.