The Forensic Metascience AI toolkit

The Forensic Metascience AI toolkit coordinates several AI agents to perform a deep forensic audit of a scientific paper. The full system checks images, text, and data in both the paper and supplementary files. Central to the system is the tools agent, which is equipped with 40+ tools for "sanity checking" the statistics and data presented in scientific papers. Many of the tools are derived from techniques described in James Heathers' book An Introduction to Forensic Metascience.

See findings and
PubPeer comments

The data integrity problem

2%

of researchers admit to fabricating or falsifying data. (Fanelli 2009 meta-analysis of surveys)

3.8%

of biomedical papers contain inappropriate image duplication. (Bik et al. 2016)

14%

of 521 trials submitted to the journal Anaesthesia contained false data. (Carlisle, 2021)

~14%

of papers contain some form of anomalous / highly questionable data (James Heathers' 2024 'highly non-systematic' review)

53.7%

of 6,200 medical residents in China admit to having committed at least one form of research misconduct. (Chen et al. 2024)

3%

of 600 biomedical datasets from well-cited publications have serious copy-paste issues (not yet proven as fraud but very suspicious). (Englund, 2026)

The rigor problem

50%

of psychology papers report a p-value that contradicts its own test statistic (Nuijten et al. 2016)

1 in 8

psychology papers contain an error that flips the significance of a conclusion (Nuijten et al. 2016)

~50%

of 71 GRIM-testable psychology papers contain an impossible mean (Brown & Heathers 2017)

63%

of meta-analyses had a data-extraction error; 37% had an error large enough to change the result (Gøtzsche et al., JAMA 2007)

The slop problem

Formulaic single-factor NHANES papers per yearNHANES = National Health and Nutrition Examination Survey

2024* is through 9 October. China/other is coded from the first-listed author affiliation in S1 Table A. Source: Suchak et al., PLOS Biology 2025.

The oversight problem

The number of scientific papers being published each year is growing exponentially with a doubling time of approximately 10 years. Meanwhile, the number of fake papermill papers being published has been estimated to be doubling every 1.5 years. AI tools make it easier than ever for fraudsters to commit research fraud.

Meanwhile, our ability to detect fraud has been constant or decreasing. A handful of largely unpaid image sleuths appear to be responsible for most detections of image fraud thus far.

The Office of Research Integrity's output has collapsed — 2025 produced the fewest misconduct findings in its 32 years of record — and NSF's Office of Inspector General stopped investigating research misconduct in early 2025, referring all allegations back to the grantee institutions.

ORI findings of research misconduct per year

Sources: HHS OGC compilation of the ORI case-summary database (1994–2005), ORI data graphs (2006–2015), and Federal Register misconduct-finding notice counts (2016–2026).

Peer review is useful, but the peer review system is under increasing strain. Studies indicate that peer reviewers do a very incomplete job when it comes to catching serious errors. Across four independent studies in biomedicine where researchers deliberately inserted errors into manuscripts, peer reviewers only caught 25–35% of them.

Where this toolkit fits in the AI for research integrity landscape

Within-paper image duplication and manipulation
Between-paper image duplication
SI/SM/Dataverse copy-paste errors
Open Source Repo dataset copy-paste errors
Tortured phrases
AI text detection
Plagiarism detection
Check stats and arithmetic
AI peer review
Proofig AI logo
ImageTwin logo
ReviewerZero logo
River Valley Technologies logoRiver Valley Technologies
Refine logo
iThenticate logo
Reviewer 3 logoReviewer 3
ScienceDetective.org logoScienceDetective.org
Forensic Metascience AI toolkit logoForensic Metascience AI toolkit

How it works

A DOI

Attempts to pull the XML, HTML, and PDF as well as supplementary information and data. If it cannot pull from an API, it provides a list of missing PDFs for a human to try to obtain.

Extraction & conversion using pdf4llm

We use a combination of tools to convert the PDF into markdown and extract tables and images — docling, pdfplumber, PyMuPDF, and Mistral OCR.

Parallel analysis

Four independent branches feed into one review.

Image analysis module

Image finding adjudication agent

An agent to screen out false positives and innocuous duplications.

Tortured phrases module

Simple module for detecting tortured phrases.

Tool-expert agent

Points the forensic toolkit at the paper's numbers. No web access.

Peer review agent

Reads the paper as a scientist would, looking for internal contradictions, scientific mistakes, and severe methodological flaws.

Join findings

Adjudicator agent

Deduplicates findings, triages out false positives, sets the final severity score for each finding, and writes the paper-level verdict.

Save findings to database

Human review

Humans review each finding in our rapid review web application. Human feedback on findings, including marking of false positives, is stored in the database to help inform improvements to the system. A human decides whether to submit to PubPeer or email the authors or an editor, and all PubPeer comments and emails are human-written, not AI-written.

The toolkit

There are 44 tools that the tools agent has access to via MCP. Each tool is Python code which implements a particular check. View detailed information on each tool on this page.

Overview

The toolkit includes a multi-agent pipeline that performs a detailed audit of a scientific paper, including the text, images, and data.

The toolkit is a “harness” for AI. The backend can be a Codex, Claude Code, or Grok Build subscription, or API calls to any agentic AI system. The “tools agent” module equips the AI with 44 tools it can use to sanity check the statistics and data presented in scientific papers. Each tool is a Python script that performs a specific type of check. Many of the techniques implemented are taken from James Heathers' book An Introduction to Forensic Metascience.

Each check applies only in a very specific set of circumstances. The GRIM tool is a simple example: if N integers in a bounded range (e.g. a 0-10 rating scale) are averaged, then only a finite set of valid averages is possible. The GRIM test checks if the reported average is in that set. We also have a couple of different tools for recalculating p-values. The SPRITE test infers possible underlying data distributions for integer data based on the allowed range, N, mean, and standard deviation. Distributions with many values at the extremes may be implausible.

Example. The paper by Ladurner et al., Journal of Neural Transmission 112: 415–428, describes a multicentre randomised placebo-controlled trial of Cerebrolysin in acute stroke. The paper reports that "16.4% of the 78 patients" in the treatment group had an adverse event. But no whole number of patients out of 78 comes to 16.4%: 12/78 = 15.4%, and 13/78 = 16.7%. However, 11/67 = 16.42%. It appears the rate was computed on the 67 completers rather than all 78 randomised patients, despite the paper's claim that side effect rates were calculated on everyone, including non-completers. This means the side effect rate is unreliable and that the true side effect rate is likely higher.

Before running any checks, the AI is forced to classify each number by type, how it was derived, and how it is used. We found that enforcing this sort of framework was important for ensuring that the LLM applies tools appropriately and checks every single checkable number reported in a paper.

To assist with human review, we have developed a “rapid review” web application. The left-hand side of the application contains a list of findings, while the right-hand side displays the PDF. Our AI system ranks every finding on a severity ladder which helps with triage. Clicking on each finding jumps the PDF to the relevant section with relevant text and numbers highlighted. Image duplication and manipulation findings are displayed via diagrams and/or videos.

To give you an idea of how thorough our system is, it consumes 1-6 million tokens and takes 25-40 minutes to run per paper. The full system costs $0.50 - $1.00 per paper to run depending on the length of the paper and the AI subscription service used. The image analysis module and copy-paste detection module are very cheap to run (pennies per paper) and can be run in isolation.

Lessons from previous initiatives

We are not the first people to have the idea of using AI to find errors in the published scientific literature.

The first major initiative along these lines was The Black Spatula project, which was launched as a crowdsourced initiative by Steve Newman in December 2024. The goal was to run the best AI model at the time (o1 preview) on thousands of papers to catch important mistakes. Several hundred people joined the Black Spatula Discord and WhatsApp group, resulting in an initial burst of activity lasting a few weeks. The project eventually petered out after a year or so. We did an “autopsy” on the Black Spatula project to see if there was anything we could learn that would inform our own project. It appears the initiative petered out due to lack of central coordination, high rates of false positives, and lack of volunteers to perform the arduous work of reviewing all the AI outputs and deciding how to act on them.

Another initiative was YesNoError, which was also founded in December 2024. The project was funded by a Solana coin/token that reached a peak market cap of $110 million. YesNoError evaluated 37,000+ papers in two months using o1, with a stated aspiration to audit 90 million. They then publicized the flagged papers, mostly without any human verification. In March 2025, science integrity expert Nick Brown found a false positive rate of 35% in a small sample of 40 papers. Shortly after, all of the results of the run disappeared from the internet and the website pivoted to focusing solely on AI research papers.

Experiments with “targeting funnels”

We are experimenting with different ways of doing screening/targeting to decide which papers to audit with the toolkit. We have entered into an unpaid collaboration with IntelliCat to test if their system could be helpful. IntelliCat has generated “Content Credibility Index” (CCI) scores for millions of papers. For a much smaller number of papers, they have also generated Data and Observations Risk Index (DORI) scores based on “domain-specific models to assess experimental data and figures for signals of manipulation or fabrication”. Initial work with our replication database indicates that a bad DORI strongly increases the chances of replication failure by 8x, so we are particularly interested in using the DORI for targeting. We are also looking at doing screening using very cheap AI APIs on thousands of open-access papers that are available in XML format to find “canary in the coal mine” issues that may be indicative of more serious issues. Another approach to targeting we are exploring is running the toolkit on papers by authors who have retractions reported by Retraction Watch.

Applications

I. Correcting the scientific record

Scientific papers are used to inform grant funding, government policy, and medical practice (e.g., via UpToDate, OpenEvidence, Cochrane Library). They increasingly inform people’s personal healthcare decisions, as patients turn to Google Scholar and AI research assistants to augment their medical decisions. Given the role of this research in decision-making, we are particularly interested in errors that substantially change a key finding, as well as data integrity issues that undermine the paper's validity. When we find major errors, we request that journals address them through retractions or errata. We submit findings on PubPeer and publicize them on our website and social media channels. A running list of what the toolkit has found — including the comments we have posted on PubPeer — is on the Findings & PubPeer comments page.

II. As a tool for pre-publication review

Once it has been validated further, we plan to open-source our codebase so anyone can run the toolkit and open-source software developers can contribute pull requests with potential improvements.

We think it is a bit ridiculous that many researchers are currently paying to use AI peer review services like reviewer3 ($11 per review) and refine.ink ($30-50 per review). As far as we can tell, the current paid services have not been externally validated to any significant degree.

A recent analysis by Paul Litvak tested two popular commercial tools against base LLMs on 10 psychology papers where he had inserted 100 errors. Only about 20% of his errors were stats/math issues that our system is particularly suited for; the bulk of the rest were issues in experimental design. He found that reviewer3 performed worst, finding only 30/100 errors, below Gemini Flash. Refine.ink found 57/100, worse than GPT 5.5 with high reasoning, which found 77/100. Litvak notes that refine.ink cost him $8.77 per error found. Interestingly, when Litvak ensembled the outputs from five systems, the overall recall jumped to 90%. We recently ran our system on the benchmark with Opus 5 and it found 49/100. We are in the process of digging deeper into how our system performed.

III. Metascientific research

The tools agent can be used in isolation to systematically research the prevalence of different error types within and across scientific fields. When James Heathers did GRIM test checks on papers in leading psychology journals, 36 out of 71 papers with testable numbers had at least one discrepancy (51%). Previous work in 2016 with the statcheck tool found that about half of psychology papers had an issue with a p-value calculation, while 12.5% had an error that actually reversed a significant conclusion. The tools agent or simplified lower-cost versions can be used to study the prevalence of errors across many different fields.