The Proof Breakers: Mythos, CryptanalysisBench, and AI's Cryptographic Turn
Anthropic's Mythos Preview found flaws in HAWK and reduced-round AES while launching CryptanalysisBench — a 191-task test of whether frontier models can do genuine cryptanalysis.
Anthropic's Frontier Red Team published a pair of results on July 28, 2026, that sit uncomfortably close to the company's Hugging Face agent disclosure from earlier in the month — but under entirely different conditions. Using Claude Mythos Preview, a model not publicly available, researchers found mathematical flaws in cryptographic algorithms themselves, not merely bugs in how programmers implemented them.
The work arrives alongside CryptanalysisBench, a 191-task benchmark built with academics at ETH Zurich, Tel Aviv University, and TU Berlin to measure whether large language models can perform genuine cryptanalysis rather than recite textbook attacks.
Attacks on HAWK and reduced-round AES
The first major finding is an improved key-recovery attack against HAWK, a post-quantum digital signature candidate in the third round of NIST's additional signatures competition. Anthropic reported that Mythos cut the effective key strength roughly in half. HAWK is not deployed in production.
The second result targets a seven-round variant of AES-128, not the full ten-round cipher used worldwide. Mythos developed a "Möbius Bridge" fingerprinting technique that improved meet-in-the-middle attacks by 200 to 800 times over prior published work.
Discovery took about 60 hours and roughly $100,000 in API cost for the HAWK attack, with substantial additional human time spent validating the AES result.
CryptanalysisBench: measuring a dangerous capability
The companion paper on arXiv evaluates five frontier models across block ciphers, hash functions, and other primitives drawn from NIST competitions. Models break 65 to 86 percent of Tier 1 targets and produce novel findings — including a key-recovery attack on the SpoC AEAD and an error in KINDI's published CCA-security proof, both described as previously unreported.
Validation bottleneck
Anthropic's post notes a growing asymmetry: Mythos discovered the AES improvement in roughly three days of autonomous work, but researchers spent nearly a month confirming correctness.
For the research community, CryptanalysisBench offers a reproducible yardstick. For the security community, the HAWK and AES results are expected outcomes of proper stress-testing — not alarms about broken TLS today, but signals about where frontier model capabilities are heading.
Sources
- Anthropic — Discovering cryptographic weaknesses with Claude (July 28, 2026)
- arXiv — CryptanalysisBench: Can LLMs do Cryptanalysis? (July 2026)
- CryptanalysisBench — Project site (July 2026)
- Schneier on Security — Measuring LLMs' Ability to Perform Cryptanalysis (July 2026)