Source · Hugging Face Forums
media Hugging Face Forums · 6d ago · 1 view

GPT-5.6 recovers 95% of hidden messages in unseen literary cryptography benchmark

A working paper by Joseph JM Walker evaluates frontier language models on an unseen multi-channel literary cryptography benchmark embedded in the physical novel "I Wrote a Book and Made a Million Dollars (I.B. Wryten)." The study finds that GPT-5.6 independently recognized the secondary communication channel and recovered approximately 95% of the embedded material on its first pass, whereas earlier models demonstrated effectively 0% ability.

media Hugging Face Forums · 7d ago · 8 views

Levent Bulut benchmarks LLM annotation reliability on objective projection

Levent Bulut published a three-study benchmark evaluating the reliability of machine-generated annotations for a Turkish narrative corpus, revealing significant discrepancies between automated raters and human judgment. The study tested six binary craft features across 120 and 100 scenes using rule-based detectors and models including Gemini 2.5 Flash, Grok, ChatGPT 5.5, and Claude Fable 5 High.