Tool-using LLM agents that rely on the Model Context Protocol (MCP) can pull from multiple sources—search results, structured records, databases—and blend them into a single answer. That makes fact-checking harder than it looks. Standard faithfulness scores pool all evidence together and ask whether a claim is supported anywhere. They do not check whether the source the answer names is actually the one backing the claim. The authors of a new paper call this failure mode cross-source conflation: a fact may be real, but attributed to the wrong record, policy document, or patient history.

ProvenanceGuard, the verification layer described in the paper, runs after an agent produces an answer and keeps source identity intact throughout the pipeline. It reads the captured MCP trace, breaks the answer into claims, finds the most relevant source for each, checks whether that source supports the claim, and compares it with the source the answer states or implies. It then issues a per-claim verdict and an overall allow-or-block decision. Blocked answers can be repaired with a RARR-style step and re-verified.

The authors tested the system on 281 real traces from a medical agent using patient records and research articles. Human experts judged 139 claims as unsupported; ProvenanceGuard caught 138 of them and let one through. It also held 67 supported claims for review or repair, reflecting a deliberately conservative policy suited to data-sensitive settings. The reported results come from a local setup using MiniLM, a DeBERTa NLI verifier, and a local language model; the authors note the approach can be adapted to hosted models, but would require its own testing and calibration.