A follow-up study tested Claude Fable 5 and ChatGPT 5.5 against a human annotator on the inferential feature of materialized metaphor using 100 scenes. The analysis revealed significant discrepancies in detection thresholds among LLMs, with Claude identifying 78 instances and ChatGPT 40, compared to near-zero counts for other models.

  • On materialized metaphors, human annotators found 9/100 instances, while Grok found 0, Gemini ~2, ChatGPT 40, and Claude 78.
  • Agreement scores (κ) were approximately zero for the hardest features, suggesting the definition itself contributes to the variance.
  • The study corrects earlier claims that atmosphere contradiction was machine-proof, showing measurable agreement for Claude (κ 0.27) and ChatGPT (κ 0.18).

The dataset and deterministic scoring script are available on Hugging Face for recomputation without model involvement.