Researchers identify "stale-document poisoning," a temporal alignment failure in Retrieval-Augmented Generation (RAG) where outdated external evidence causes models to provide incorrect answers despite knowing the right response without retrieval.
- The study constructs a benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy.
- Outdated retrieval flips 30% of Llama and 37% of Qwen answers even without instructions to trust the document.
- Explicit follow instructions raise poisoning rates to 66% for Llama and 75% for Qwen.
- Poisoning ranges from 17-91% across four open models and domains, while up-to-date evidence is followed in 97-100% of trials.
- A fixed recency-aware hybrid re-ranker reduces poisoning by 4.6-10.0 points when dates are accurate.
The findings indicate that reliable RAG requires selective trust, where models must determine not only what retrieved evidence says but whether it still applies.