Levent Bulut has updated the operational definition of "summarization bias" from a loose description to a testable construct, establishing a registered research protocol with pre-specified falsifiers. The new definition characterizes the bias as the systematic tendency of models to replace shown-mode structure with abstract summary labels (told mode), rather than merely stripping details.
- The construct is defined as replacing reconstructable shown-mode structure with the abstract summary label that names its content, specifically favoring "told" declarations over "shown" physical cues.
- The protocol requires 20 matched pairs of human-written stimuli and evaluation by at least three model families using LLM judges.
- A key falsifier is established: if models show no preference for told mode compared to humans (Cohen’s h ≥ 0.3 or OR ≥ 2), the construct will be withdrawn.
- Discriminant checks are included to ensure the bias is dissociated from verbosity and sycophancy by re-running tests with length-mismatched pairs and varied framing.
Bulut seeks criticism of the design, specifically regarding whether length-matching sufficiently separates this from verbosity bias and if the Suppressed Information Index is the correct instrument for measuring inferential load.