A study titled "Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation" evaluates how large language models handle false premises injected into conversation history, a failure mode termed session-level contamination. The researchers tested GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains using 22,500 turns at temperature zero.
- GPT-5.4 Mini demonstrated a content-independent policy with zero adoptions of false premises across all 500 sessions.
- Gemini-3.1 Flash-Lite followed a steep authority gradient, showing adoption rates ranging from 0.1% for self-attributed falsehoods to 94.0% under instruction override.
- GLM-4.5-Air exhibited a shallower gradient (15.8% vs 84.2%) and recovered in 94.5% of affected sessions, compared to only 60.0% for Gemini.
The authors conclude that conversation history acts as an untrusted attack surface requiring provenance-aware system design, and they have released the complete framework as an open-source benchmark.