Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · September 30, 2026 · concept

ProvenanceGuardMultiverse ComputingMCPmcpagent-securityrag

MCP Agents Can Credit a True Claim to the Wrong Source, and Your Fact-Checker Will Pass It

ProvenanceGuard shows why multi-server agents need per-source verification, what it costs to add, and where it still guesses

"According to the account record, this plan includes a 30-day refund window."

Every word of that sentence can be true and it can still be wrong. The refund window exists. It lives in the policy document. The account record says nothing about it. An agent that pulled from both sources blended them, and the answer now points a reader at the wrong system as its authority.

That example comes from a September 29 post on the Hugging Face blog by three researchers at Multiverse Computing, and it names a failure most agent eval stacks cannot see. They call it cross-source conflation. Their tool for catching it, ProvenanceGuard, is the most practical thing I've read this week about MCP, because it goes after the part of an answer nobody tests: the citation.

Why pooled fact-checking misses this

Most groundedness checks work the same way. Take the answer, take all the retrieved evidence, and ask whether the evidence supports each claim. MiniCheck does a version of this. So does RAGAS Faithfulness. For a single retrieval source that's fine, because there's only one place the fact could have come from.

MCP changes the shape of the problem. A typical agent now talks to several servers at once: a CRM, a docs search, a ticketing system, a database. Their outputs land in one context window. When the checker pools that evidence, it asks "is this claim supported somewhere?" and the answer is yes. The authors put it bluntly: "A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and passes it."

So the fact passes. The attribution was never checked.

Think about where that bites. In a support flow, "per your account" versus "per our policy" decides whether a customer believes they have a contractual right. In a finance agent, a figure credited to the wrong filing sends an analyst to the wrong document. The post's clinical example is the sharpest: a medication detail pulled from a patient-history tool "becomes misleading the moment the answer presents it as a finding from the medical literature." The words match. The authority behind them doesn't.

And the more MCP servers you wire in, the more places a true fact can be misfiled.

How ProvenanceGuard checks the citation

The design is a pipeline of five small steps, run after the agent writes its answer:

  1. A local language model splits the answer into individual claims.
  2. A MiniLM embedding model routes each claim to the source it most likely came from.
  3. A DeBERTa-v3-base NLI model checks whether that source supports the claim.
  4. An attribution check compares the source that supports the claim with the source the answer names or implies.
  5. A calibrated decision step decides whether to pass the claim or hold it for review or repair.

The attribution check is the whole point. Without it, the pipeline is a decent groundedness checker. With it, the pipeline becomes a provenance checker, because it asks a second question: supported by which source, and is that the one you told the reader?

The models are small on purpose. The post reports roughly half a second per answer, with the routing and NLI calls themselves in the tens of milliseconds. That's cheap enough to run on every response in production, not just in an offline eval.

What the numbers say, and what they don't

The evaluation used 281 real medical agent traces. From those, the team held out 40 answers containing 361 claims and had experts label them.

Experts said 139 of those claims should not pass. ProvenanceGuard caught 138 and let one through. On the block decision it scored 0.802 F1, ahead of MiniCheck at 0.783, RAGAS Faithfulness at 0.758, AlignScore at 0.662 and SummaC-ZS at 0.436.

That margin over MiniCheck is small, and I don't think it's the headline. The headline is the swap test. The team took answers and changed only the named source, leaving the supporting evidence in place. A pooled checker should pass every one of those, because the fact is still supported somewhere. ProvenanceGuard caught all 50 swaps.

That's the capability that didn't exist in the baselines.

Now the costs, in the authors' own numbers. To catch 138 of 139, it also held 67 claims experts considered supported and sent them for review or repair. On a 361-claim set, that's a lot of friction, and the post is explicit that the cautious setting "favors a second look at some supported claims over letting unsupported ones through."

The harder number is source identification. For claims with an identifiable source in the main test, it named the right source about 86% of the time. On a separate test built with several similar sources, that fell to 50.3%. Its block decision held up there (0.846 F1), so it still knew something was off. It just couldn't reliably say which source was the real one. The authors name this as their open problem: "Telling similar sources apart remains an important area for improvement."

Keep that in mind when you read "caught all 50 swaps." A swap between a billing database and a clinical guideline is easy to spot. A swap between two versions of the same policy document is the case your agent will actually produce.

The integration that already shipped

The post says NVIDIA's NVFlow merged an optional grounding-verification stage for its finance agent. It checks completed answers against the SEC excerpts the agent retrieved and writes its decisions to a separate results file, without changing the original rollout or training data.

The pull request itself is worth reading for its framing. The author describes a post-rollout GroundingVerifier that checks claims about entity, metric, date, value and filing attribution, runs on CPU, and uses deterministic claim decomposition with MiniLM routing and DeBERTa NLI. The PR reports 99.17% acceptance of legitimate answers (119 of 120 eligible traces) and rejection of all 120 simulated adversarial attacks, from numeric fabrication to entity, metric and date conflation. That's a different setup from the blog's medical evaluation, with synthetic attacks and a small sample, so treat it as the author's report rather than a second benchmark.

The design choice that matters: it runs beside the pipeline, not inside it. That's the pattern I'd copy.

Put this into practice

You don't need ProvenanceGuard itself to start. You need the record it depends on.

Log which tool call produced each piece of evidence. Every MCP response your agent receives should carry a server name and call ID into whatever you store. If your traces only keep the final context window, you've already lost the information a provenance check needs.

Make the agent name its sources in a structured way. Free-text "according to the account record" is hard to verify. A citation field that references a tool-call ID turns the attribution check into a lookup instead of a guess.

Run the swap test on your own agent. Take 20 answers, change only the named source, and run them through whatever groundedness check you use today. If it passes most of them, you now know your checker is source-blind. This costs an afternoon and no new tooling.

Add a per-source NLI check off the hot path first. Copy the NVFlow pattern: verify completed answers and write decisions to a separate log. Watch the hold rate for a week before you let it block anything.

Separate look-alike sources at design time. If two MCP servers return near-identical content (staging and prod docs, two policy versions), that's exactly where source identification collapses. Merge them, tag them with version metadata, or drop one.

Honest limitations

This is one team's evaluation on one domain. The main results come from 40 held-out answers drawn from medical traces, and I could not find a statement in the post about whether that data is available for others to rerun.

The 67 held-but-supported claims are not a rounding error. In a customer-facing flow, a checker that holds that many good answers will get turned off by someone under deadline pressure. Budget for a review queue or tune the threshold before you ship it.

A 50.3% exact-source rate on similar sources means provenance checking is not yet a way to assign attribution. It's a way to flag that attribution needs a human.

The approach also assumes the agent's answer names or implies a source. An agent that never cites anything sails straight past the attribution check. Provenance checking only works if you first require provenance.

And I couldn't read the arXiv paper behind the post during this research (the fetch was rate-limited), so everything here comes from the blog and the NVFlow pull request.

Where this leaves your agent

For two years the question about agent answers was "is it true?" With several MCP servers feeding one context, that question stopped being enough. A true fact under the wrong name is still a wrong answer, and it's the kind your current checks will wave through.

You can find out this week whether your stack sees it. Swap the sources on a handful of answers and watch what your checker does. If it shrugs, you know where your next week of work goes.

Sources: Multiverse Computing, "Getting the Source Right, Not Just the Fact," Hugging Face blog, September 29, 2026; NVIDIA NVFlow pull request #9; MiniLM and DeBERTa-v3-base NLI model cards.


Medium metadata

  • Title: MCP Agents Can Credit a True Claim to the Wrong Source, and Your Fact-Checker Will Pass It
  • Subtitle: ProvenanceGuard shows why multi-server agents need per-source verification, what it costs to add, and where it still guesses
  • Tags: MCP, AI Agents, LLM Evaluation, RAG, AI Safety
  • Canonical URL: fervorai.dev (import from the published article URL)
  • Reading time: about 8 minutes