September 29, 2026 — Hugging Face published Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents, a Multiverse Computing post on ProvenanceGuard. The paper is ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (arXiv 2606.18037).
The hook for anyone shipping Model Context Protocol agents is simple. Your faithfulness eval can go green while the citation is still wrong. explainx.ai's read: supported somewhere is not a release gate. Supported by the source the answer named is.
That is a different class of error than indirect prompt injection or a malicious server. It is a tool-result handling bug: you mixed evidence, then scored the soup.
TL;DR — what should an MCP builder change?
| Question | Direct answer |
|---|---|
| What shipped? | Hugging Face blog (Sep 29, 2026) + paper arXiv:2606.18037 |
| Who wrote it? | Multiverse Computing (blog: Antonio Tiene, Ander Alvarez Sanz, Oliver Wirjadi; paper: Ander Alvarez, Santhiya Rajan, Samuel Mugel, Román Orús) |
| Failure mode | Cross-source conflation — true fact, wrong source |
| What ProvenanceGuard is | Post-generation layer on a black-box MCP agent; no retraining |
| Held-out block F1 | 0.802 vs MiniCheck 0.783, RAGAS Faithfulness 0.758 |
| Source accuracy (eligible claims) | 0.858 over 260 source-eligible claims |
| Expert miss | Experts wanted 139 claims blocked; the system caught 138, let one through |
| Conservative cost | Held 67 claims experts treated as supported (review/repair) |
| Attribution swap probe | 50/50 injected source swaps detected |
| Repair | All 173 blocked full-trace answers resolved; 144 ended in fallback, not a rewrite |
| Harder multi-source slice | Block F1 0.846; exact source 50.3%; source-plus-relation 0.229 |
| Latency (reported local setup) | Roughly 0.5 s per answer |
| Your eval question | Does “grounded” mean any tool output, or the named source? |
Primary: Hugging Face blog · arXiv:2606.18037.
What people are asking after the headline
Is this just hallucination with extra steps?
No. Hallucination is a claim with no support. Cross-source conflation is a claim that is supported — in the wrong drawer.
The Hugging Face post’s customer-support example: “According to the account record, this plan includes a 30-day refund window.” The 30-day window can be real in a policy document. Pool the account record and the policy and a RAG-style faithfulness metric still sees support. Keep sources separate and the attribution is false. In a data-sensitive product, wrong provenance can be as damaging as a wrong number.
The clinical analog in the same post: a medication detail from a patient-history tool, presented as a finding from the medical literature. The fact is not invented. The source family is.
Why does MCP make this worse than single-passage RAG?
Because MCP is built for heterogeneous tools: search, account APIs, databases, records, metadata. The post’s opening is explicit: agents no longer read from a single retrieved passage. They call several servers and weave one answer.
Classic checkers named in the blog — RAGAS faithfulness, MiniCheck, AlignScore, SummaC — ask whether a claim is supported once evidence has been pooled. In their usual form they do not say which MCP tool output supports the claim, or whether that is the source the answer names.
If your harness logs tools.concat(results) into one evidence string, you have already thrown away the object the verifier needs.
Do I need their MiniLM + DeBERTa stack?
Not to adopt the contract. The blog is careful: MiniLM for routing, DeBERTa NLI for support, a local LM for claim split, and a calibrated decision step are the evaluated setup, not a requirement of the method. Hosted models can run the same claim / source / decision steps; they need their own calibration.
The paper’s retained calibrator is a random forest (400 trees, max depth 5) on verifier-internal features, operating threshold 0.65, selected on validation by reject/block F1. That is plumbing. The product requirement is: never collapse source identity.
What ProvenanceGuard actually checks

The Hugging Face post lists five sequential steps after the agent answers:
- Break the answer into specific claims.
- Find the source most relevant to each claim.
- Check whether that source supports the claim.
- Compare that source with the one the answer names or implies.
- Emit a per-claim source verdict and a global allow or block.
The paper’s trace object is (tool_i, source_i, text_i). Multiple calls to the same tool type with different source IDs stay separate evidence objects. If a tool output lacks a stable source ID, the verifier falls back to tool name — which is already better than an anonymous blob, and already worse than a real document ID.
Support is not “the sentence sounds right.” Literal values — numbers, dates, identifiers — cannot pass if they are absent from the routed source. Neutral NLI is not enough evidence unless a narrow lexical rescue applies on structured fields.
Attribution is a second decision. The paper is explicit: calibration answers only whether the claim is supported by the routed source. It does not yet decide whether the answer described that source correctly. Conflation is marked when support is positive but the stated family and the routed family are incompatible.
Ander Alvarez, responding on the Hugging Face thread, restated the product definition: when a claim is “grounded,” it means it is supported by the source the answer names or implies. Finding the fact somewhere else in the tool outputs is not enough.
That is the sentence to paste into your eval spec.
Numbers you can cite without inflating them
The empirical testbed is a medical MCP agent (patient records, research articles, other tools) — 281 captured traces. Medicine is useful because a fact from a patient record and a fact from general research cannot share a source family. The method is not limited to medicine; it assumes the trace keeps tool outputs and source IDs.
Held-out scoring used 40 answers / 361 claims, with human expert checks on those claims. A larger 266-trace subset produced 2,325 LLM-assisted labels for train/val/test splits by trace (so claims from one answer do not leak across splits).
| Verifier | Reject/block F1 | Emits claim-to-source ID |
|---|---|---|
| ProvenanceGuard | 0.802 | Yes |
| MiniCheck | 0.783 | No |
| RAGAS Faithfulness | 0.758 | No |
| AlignScore | 0.662 | No |
| SummaC-ZS | 0.436 | No |
Source: Hugging Face results table (paper abstract and Table III). On source-eligible claims, source accuracy is 0.858 (260 claims).
The blog’s most concrete safety line: experts said 139 claims should not pass; ProvenanceGuard caught 138 and let one through. It also held 67 claims the experts considered supported. Fail-closed on purpose: second look beats a silent pass.
Do not mix metrics. On a harder multi-source test, block F1 is 0.846, but exact source identification is 50.3%, and the paper’s source-plus-relation accuracy drops to 0.229. Blocking unsupported claims and naming the precise neighboring source are different jobs. Similar sources remain hard.
The 50 controlled attribution swaps (supporting evidence left intact, named source changed) were all caught. That probe shows the system can see a clear source error. It does not prove you can always pick among many near-duplicate documents.
Repair (RARR-style revise and re-verify): all 173 blocked answers on the full-trace run resolved; 144 became fallback text rather than a substantive rewrite. On reconstructed multi-source traces, all 59 initially blocked answers resolved with two terminal fallbacks. Offline overhead on the reported local configuration is about half a second per answer; NLI and routing themselves are tens of milliseconds.
The paper’s central claim is limited: source-attribution factuality in MCP-grounded answers. It does not claim to solve open-domain factuality, domain safety validation, or parametric-knowledge correction. Hugging Face quotes the same narrow goal: not proving a source is clinically or scientifically correct.
What to change in tool-result handling
This is the practitioner section. Pair it with the permission and injection controls in the MCP security guide. ProvenanceGuard does not replace least privilege. It replaces pooled faithfulness as the only truth check.
1. Stop flattening MCP results before they hit logs or judges
If the host concatenates every tools/call payload into one context field, source-aware verification cannot run. Your agent harness should persist an array of evidence objects, not a string.
type McpEvidence = {
tool: string; // MCP tool name
sourceId: string; // document, record, URL, or row key — not "search"
callId: string; // this invocation, even if tool name repeats
text: string; // raw tool output, not a summary you cannot re-check
};
// Wrong: one blob for RAGAS / the next LLM judge
const pooled = results.map((r) => r.text).join('\n\n');
// Right: keep identity through the answer and the eval
const evidence: McpEvidence[] = results.map((r) => ({
tool: r.tool,
sourceId: r.sourceId ?? r.tool,
callId: r.callId,
text: r.text,
}));
If the server does not return a source ID, mint one at the client from URL, primary key, or content hash, and store it next to the bytes the model saw. Fallback-to-tool-name is the paper’s last resort, not a design goal.
2. Treat “according to …” as a checkable claim, not flavor text
If the model writes “according to the account record,” that span is an attribution. Your decomposer should keep it. A source-blind judge that only looks at the refund-window clause will miss the bug the Hugging Face example is built around.
Implicit attribution counts too. “The chart shows …” when the number came from a literature abstract is still conflation. If you cannot map the phrase to a source family, mark the claim unavailable rather than guessing — the paper does that for ambiguous cases.
3. Split support scoring from attribution scoring
Two booleans, not one:
- Supported by routed source? NLI / overlap / protected literals against one evidence object.
- Named source matches routed source? Alias match on record vs policy vs search vs database.
Allow the answer only if every factual claim passes both. That is the paper’s fail-closed answer policy: conflation, contradiction, missing protected values, failed routing, or insufficient support all block.
4. Do not let the judge see the whole soup while deciding one claim
ProvenanceGuard routes first (centroid of MiniLM chunk embeddings, cosine to the claim), then narrows the premise, then runs NLI. Premise length in the reported service default is 512 tokens for the pair — enough to avoid a 256-token bottleneck, still inside the DeBERTa window they used.
If your “verifier” is a second LLM call over trace + answer, you can still prompt it to score one claim against one sourceId. The Hugging Face comments ask whether a second LLM replaces the random forest. The paper does not publish that ablation. Until you have one, do not assume a chat completion is calibrated the same way. What you can copy immediately is not pooling.
5. Repair is a product decision, not a license to invent
When blocked, the reported loop can rewrite, prune, or emit fallback with no evidence-requiring factual claim. On real traces, most resolutions were fallback. That is the conservative product: say less rather than cite the wrong drawer.
Wire this next to human review, not instead of it. Human-in-the-loop still owns the residual one miss on 139 should-block claims.
6. Protect the trace itself
Source-aware verification assumes the trace is honest. If the agent can rewrite tool logs — the failure METR documented as tool-call spoofing — a post-hoc verifier reads fiction. Sign or hash tool results at the host, outside the model’s write path. Store traces off the git remote the way PixelLeak hygiene stores screenshots: they are sensitive artifacts, not chat candy.
Hugging Face’s own security.txt note for agents is a different incident class (redirect curious agents). The shared discipline is the same: write down what the agent is allowed to treat as evidence, in a place the harness actually loads.
How this sits next to MCP security and RAG
The MCP security guide is still the right checklist for auth, confused deputy, injection in tool descriptions, and audit logs. ProvenanceGuard adds a content axis those controls do not cover: the agent used only allowed tools and still mis-cited them.
RAG vs MCP comparisons on explainx.ai already argue that MCP is live tools, not a static index. This paper is the missing eval: once you have live tools, citation quality is tool-provenance quality. ALCE-style citation checks on a retrieved set are the closest prior task; MCP adds stable tool-level IDs and a routing step the authors say passage-level citation evals do not provide.
The blog notes NVIDIA NVFlow merged an optional grounding-verification stage for a finance agent against retrieved SEC excerpts, using the source-aware approach; the RARR repair loop belongs to the broader research system, not that product merge. Treat that as an existence proof that post-generation, freeze-the-rollout verification is shipping in at least one vendor stack — not as a requirement to buy anything.
ProvenanceGuard was also a poster at the Agentic AI Summit 2026 at UC Berkeley. That does not change the implementation bar: capture traces with IDs, or you cannot reproduce the paper’s interface.
Honest limitations
- Medical traces, not prevalence. The 281-trace corpus is not, by the authors’ own statement, an estimate of how often conflation happens in the wild. Random captured traces were often single-family. The 50 swaps are a constructed probe.
- Conservative false blocks. Sixty-seven expert-supported claims still went to review or repair on the held-out packet. Latency-sensitive chat UX will feel that.
- Neighboring sources. Block F1 can stay high while exact ownership collapses when candidates are semantically close. Do not advertise 86% source accuracy as the multi-source number.
- Local models in the paper. Reported F1 is for MiniLM / DeBERTa / local LM plus the forest. A hosted judge needs a new calibration study.
- Not clinical validation. Synthetic or de-identified benchmark artifacts; not a medical-device study.
- Verifier compromise. A plugin that reads traces is another trust boundary. Scope it like any other MCP server: least privilege, no extra egress, logs of its own verdicts.
- No explainx.ai re-run. We did not execute ProvenanceGuard. Figures are from the September 29, 2026 Hugging Face post and arXiv 2606.18037 as fetched for this article.
Checklist you can run this week
| Control | Owner | Done when |
|---|---|---|
Evidence array with tool, sourceId, callId, raw text | Harness | Traces in staging match the paper’s e_i shape |
No pooled evidence string for production judges | Eval | RAGAS/MiniCheck (if kept) run per source, or are labeled source-blind |
| “According to X” parsed as attribution | Product | Conflation fixtures fail the release gate |
| Protected literals (IDs, dates, doses, prices) must appear in the named source | Quality | One fixture per tool family |
| Fail closed + fallback copy, not silent pass | Product | Blocked answers never ship a naked wrong citation |
| Trace integrity (hash/sign at host) | Security | Model cannot overwrite tool results |
| Human review queue for blocks | Ops | Residual miss on should-block claims has an owner |
If you only do one row, stop concatenating tool results. Everything else in ProvenanceGuard is details on top of that.
Related on explainx.ai
- MCP Security Guide 2026
- What is MCP? Model Context Protocol complete guide
- RAG vs MCP
- What is indirect prompt injection?
- Hugging Face security.txt note for AI agents
- PixelLeak: agent screenshots on public GitHub
- OpenAI agents spoofed tool calls to trick evaluators
- What is an agent harness?
- Human-in-the-loop AI
Primary sources: Hugging Face — Getting the Source Right, Not Just the Fact (published September 29, 2026) · Alvarez, Rajan, Mugel, and Orús, ProvenanceGuard, arXiv:2606.18037.
Figures, quotes, and setup details in this post follow the Hugging Face blog dated September 29, 2026 and arXiv 2606.18037 as retrieved for publication. Model identifiers, F1 scores, and repair counts may change in later paper revisions. This article does not include exploit steps or instructions for attacking MCP servers.
