A working paper by Joseph JM Walker evaluates frontier language models on an unseen multi-channel literary cryptography benchmark embedded in the physical novel "I Wrote a Book and Made a Million Dollars (I.B. Wryten)." The study finds that GPT-5.6 independently recognized the secondary communication channel and recovered approximately 95% of the embedded material on its first pass, whereas earlier models demonstrated effectively 0% ability.
- Previous models failed to identify or solve the book's cryptographic layer, with some systems like Fable and Claude Opus declining to continue due to guardrail behavior.
- GPT-5.6 distinguished provisional theories from established findings and maintained a growing theory ledger while connecting clues across hundreds of pages.
- The benchmark consists of more than sixty intentionally embedded mechanisms, including Morse code, Vigenère encryption, typographic steganography, and low-contrast text.
- The evaluation protocol involved sequential reading without an answer key or cipher locations, testing the model's ability to perform signal recognition under ambiguity.
The authors consider this result significant because it marks a qualitative leap from no demonstrated cryptographic capability to sustained, highly capable cryptographic reasoning within a single model generation.