A researcher is requesting a single independent human annotator to label 100 Turkish narrative scenes in order to determine whether low inter-rater agreement stems from the interpretive nature of the task or from underspecified annotation definitions. The study found that four LLMs and a rule-based detector agreed with each other and the human reference at roughly chance level, with Cohen’s κ scores ranging from 0.000 to 0.185.

  • The corpus consists of 100 bilingual Turkish–English scenes annotated for six craft features, including materialized metaphor.
  • Five machine raters scored against a locked set of human reference labels containing 9 positives out of 100.
  • Raw agreement across the five raters was 74.7%–86.3%, but this is misleading due to extreme class distribution and the kappa paradox.
  • The researcher distinguishes between two possible explanations: the feature is inherently interpretive, or the definition is clear to the author but underdetermined for others.

The result will be published regardless of the outcome, providing a substantive finding on whether such features can be delegated to machines or require clearer human guidelines.