← The retracted research file, every correction kept inline
AVE (the Agentic Vulnerability Enumeration standard) keeps a research directory documenting every pass we've run to check the corpus against the current threat landscape. One of those files, a June 2026 benchmark report mapping five external MCP security datasets against AVE's coverage, turned out to be substantially fabricated: invented per-class taxonomies attributed to five real papers, one outright fabricated paper title, and every downstream coverage number computed from those invented lists. This is what we found, how we found it, and what we did and didn't do about it.
What was wrong
The file claimed to map five published research datasets — MCPSecBench, a formal MCP threat-model paper, Hou et al.'s MCP taxonomy, MCP-SafetyBench, and MCPTox — against AVE's own record set, class by class, to find genuine coverage gaps. An audit checking every one of those class-level claims against each paper's real, current text found:
102 checkable claims. 30 confirmed correct. 72 confirmed wrong. 0 unverifiable.
The shape of the failure matters more than the count. This was not miscitation — a right paper with a wrong page number, a rounded figure, a class renamed since drafting. Every cited paper is real. The dataset-level facts, mostly just the class counts, were largely accurate: MCPSecBench really does define 17 attack types, MCP-SafetyBench really does define 20. What's wrong is everything under that: the file listed specific class names for each dataset, and for four of the five datasets, the large majority of those names do not appear anywhere in the paper they're attributed to. Tool Poisoning, Prompt Injection, and a handful of others matched. Credential Theft, Remote Code Execution, Privilege Escalation, Server Impersonation, Cross-Server Contamination, and a dozen more, all attributed to MCPSecBench specifically, do not exist in that paper's real 17-class taxonomy at all. The fifth dataset, MCPTox, was wrong at a deeper level than any individual class name: the file described it as a content-safety and toxicity benchmark. The real paper is entirely about tool-poisoning attack detection, evaluated against 45 real MCP servers and 353 real tools. It has no content-safety dimension whatsoever. None of the eleven class names attributed to it exist in the real paper, because the real paper isn't about what the file said it was about.
One more find, smaller but worth naming precisely: one of the five papers was cited under a title that isn't its real title. The paper exists, the arXiv ID resolves, and it says something different from what its invented title implied.
Because every per-dataset "AVE coverage" table was computed against these invented class lists, none of that analysis holds either. The file's headline conclusion — that AVE had exactly one confirmed genuine coverage gap, a missing resource-exhaustion / agentic-DoS record — named a class that does not exist in either of the two papers cited as independently confirming it.
How it was caught
Not by a dedicated review. A roadmap item — "add the resource-exhaustion record this benchmark identified as our one confirmed gap" — sat unresolved across two separate revisions of an internal tracking document, each time deferred rather than actioned. On the third pass, instead of carrying it forward again, it finally got checked directly against the two papers it cited. Neither contained the class. That produced a small, boring-looking GitHub issue: two specific citations, checked and shown wrong.
The decision that mattered wasn't finding those two. It was refusing to treat two wrong citations as the finding. Two wrong citations in a document making dozens of citation-shaped claims is a sample, not a conclusion, and the only way to know whether it generalizes is to check the rest — so the entire file was audited, table by table, class by class, against each paper's actual current primary text rather than against a summary or against the file's own internal consistency.
What we did about it
The file was not deleted, and its wrong tables were not silently corrected in place. Every retracted claim is still visible in the document, marked inline, next to what checking it actually found — the paper's real class list, quoted directly from its own tables, sitting right beside the invented one it replaces. This is the same practice we apply to every negative result: publish it, don't edit it away, because a corrected document with no visible trace of what was wrong is much easier to trust today and much easier to repeat tomorrow.
What we explicitly did not do is re-derive the coverage analysis. Figuring out what AVE's real relationship is to each paper's real, actual taxonomy is a substantive research task — five datasets, real class-by-class mechanism comparison — not a citation check, and improvising it inside the same pass that just found the fabrication would risk repeating the exact failure this pass exists to correct. The file says this plainly, dataset by dataset: coverage against the real taxonomy has not been re-derived. That's a separate, future piece of work, not something we're claiming credit for having already done.
Containment
A wrong number that stays wrong only inside the file that produced it is a research-quality problem. A wrong number that reaches somewhere else — another document, a public comment, an external thread — is a different problem, and one that doesn't wait for a convenient time to fix.
We checked. Every retracted figure and claim was grepped across the full repository, and every GitHub issue and pull request on the project's own repo was searched for the same terms. Two live, uncorrected instances turned up, both inside the project's own changelog: a release-notes entry that had credited three shipped records to this benchmark's "confirmed genuine gaps" analysis, and a stale roadmap bullet that still asserted the resource-exhaustion gap as real after it had already been dropped elsewhere. Both are corrected now, in place, with the same visible-not-silent convention as the source document itself. Neither correction touches the three records in question — each of them carries its own independent, verified primary sourcing (an RFC, a CWE entry, an OWASP document) that has nothing to do with the retracted analysis; what was wrong was only the sentence crediting their discovery to it.
Beyond the repo, we checked every external thread this project is currently party to that could plausibly carry a number sourced from this file — a pilot mapping submitted to a related OWASP project, an open numbering discussion on another OWASP project's tracker, two threads with an adjacent standards body, and the project's own prior public issue threads — read in full, not just grepped. And we checked for any social-media post that might have cited these figures.
We found nothing in any of them. That's a real result, not a formality, and worth stating as plainly as the fabrication itself: this stayed contained to the one document that produced it.
What it touched, and what it didn't
We checked whether any of the corpus's 80 published records is actually sourced from this file. None is. The one new record the file's own "genuine gap" analysis recommended was never created — its target ID was later reused for an unrelated, independently-sourced record. And a direct search of every published record's own text for this file's name, or any of its five dataset names, returns zero matches. Whatever this file got wrong, it did not make it into anything the corpus asks anyone to rely on.
Checking the rest of the shelf
A citation practice that produces one bad document and nowhere else is a different, better situation than a practice that produces bad documents whenever nobody happens to check. So we checked the other file living alongside this one in the same research directory — a separate benchmark pass, five CVEs, one peer-reviewed paper, a vendor disclosure report, a research-body corroboration note, a MITRE ATLAS technique ID, and a news report, eleven checkable external claims in total.
All eleven confirmed correct against their real, primary sources. Zero wrong. Zero unverifiable.
That result matters as much as the 72 that were wrong in the other file. It means this wasn't a standing practice of citing without checking — it was one document, produced one way, that didn't get the same discipline applied to it before it was committed. That's a real, meaningfully different situation from a practice-wide failure, and it's worth saying so exactly that plainly rather than either overclaiming systemic rot or quietly hoping the one bad file was the whole story.
Why this is the post, not a footnote
It would be easy to write a piece about verification discipline that leads with someone else's mistake — and there are real ones to point to. But publishing that framing while sitting on a worse, still-uncorrected failure of our own would be exactly the dishonesty a post about verification discipline exists to argue against. The actual news here is not that citation-checking catches errors in general. It's that it caught this one, inside our own published research, after it had already sat unfixed for two review cycles — and that finding it changed what came next: audit the whole file rather than patch two lines, contain the damage before writing anything else, check the sibling document rather than assume it was fine too, and only then write this.
None of that is a success story about a well-run verification process. The process took months to catch this, and it was caught late, by refusing to let a small finding stay small — not by the process working as designed the first time. What we can actually stand behind is narrower and more honest: once caught, it was checked as thoroughly as we know how, corrected visibly, contained rather than left to spread, and the parts we haven't finished — re-deriving real coverage against each paper's actual taxonomy — are named as unfinished rather than quietly assumed done.
The retracted file, in full: docs/agents/research/benchmark-2026-06.md, every fabricated table kept visible next to its correction. The audit and containment work: issue #241 and PR #249.