ProvenanceGuard is a post-generation verification layer for Model Context Protocol (MCP) agents that detects cross-source conflation by ensuring claims are attributed to the correct evidence source. The system analyzes captured MCP traces without retraining the agent, breaking answers into claims and verifying that the supporting source matches the one named or implied in the output.
- ProvenanceGuard caught 138 of 139 claims that human experts deemed unsupported while correctly identifying the right source about 86% of the time.
- In a test with similar sources, it scored 0.846 F1 for blocking decisions and identified exact sources in 50.3% of cases.
- It successfully detected all 50 instances of wrong attribution where the named source was swapped while evidence remained intact.
- The system integrates a RARR-style repair loop that resolved all blocked answers, though many resulted in safe fallback text rather than substantive rewrites.
This approach addresses the limitation of faithfulness scores by making provenance visible claim-by-claim, which is critical for data-sensitive settings where wrong attribution can be as damaging as incorrect facts.