The toolkit
This page documents 46 tools in the Forensic Metascience AI toolkit. Every one is a deterministic function, not a language model: it takes numbers the paper already printed, does arithmetic, and returns a structured result with a severity. The tools agent chooses which tools to point at a paper and how to read what comes back — it never decides what a check concludes.
None of these need the raw data. That is the point of the field: a paper's own summary statistics constrain each other tightly enough that fabrication and error leave traces in the numbers as published.
The tools agent’s registry holds 58 tools. The ones not listed here are left out on purpose, and for three different reasons. Ten read PDFs, publisher XML, HTML and supplementary files into checkable numbers — that is how the figures are obtained, not a check on them. The image-integrity screen is a pipeline of twelve discrete stages rather than a single check, so it has its own page. And the tortured-phrases screen is withheld while it remains quarantined behind a seed dictionary — it runs, but it is not yet calibrated enough to describe as a working check.
Each card below states the worst verdict its tool can deliver. That ceiling is a real limit, not a formality — most of these tools refuse rather than guess when a premise they depend on is missing, and several are capped below what their arithmetic would technically license because an innocent explanation always remains. Two are marked quarantined: they run, but the pipeline deliberately withholds their verdicts until their calibration criteria are met.
Granularity — the GRIM family10 tools
Integer data cannot produce just any summary statistic. Whole-number responses — Likert items, counts, binary outcomes — force means, percentages and standard deviations onto a discrete lattice fixed by the sample size and the reported precision. A value off that lattice was not computed from the data described. These are the toolkit's strongest checks: they need nothing but the numbers already printed, and a failure is arithmetic rather than opinion.
GRIM
grim_checkWhen it applies
A reported mean of data where every subject contributed a whole number, with the sample size known.
How it works
The mean of n integers must be some integer total divided by n. GRIM asks whether any integer total rounds to the mean exactly as printed. The tolerance scales with n rather than being a fixed half-unit, and both the round-half and truncation reporting conventions are tried before anything is flagged.
Inputs
- mean
- numberAs printed, not recomputed.
- n
- integer
- decimals
- integerPrinted decimals of the mean — “5.9” is 1, “5.90” is 2. There is no default: a wrong value accuses, so a missing one is refused.
- scale_min / scale_maxoptional
- integers
- data_typeoptional
- integer / continuous / unknownContinuous data is refused outright rather than judged.
- discreteness_sourceoptional
- stated / logically_necessary / assumedWho says each subject contributed a whole number. On an assumed premise a granularity failure caps at suspicious.
Output
Consistent, or a failure naming its mode — granularity or range violation, reported separately rather than fused. Only a failure resting on a stated or logically necessary premise reaches impossible; assumed bounds, assumed discreteness and truncation-only passes all cap at suspicious. Values read off a figure are refused, not judged.
Reference
Brown, N. J. L., & Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. Social Psychological and Personality Science 8(4), 363–369. 2017.
GRIM for percentages
grim_percentageWhen it applies
A percentage that is a count's share of the sample — considerably more powerful than GRIM on a mean, because percentages carry more precision.
How it works
A share of a sample must be 100·k/n for some integer k, so at a given n only a finite set of percentages exists. Both round-half and truncation conventions are tried: SPSS and several journal styles truncate, and 9/93 = 9.6774 printed as “9.6%” is a truncating reporter rather than an impossibility.
Inputs
- percentage
- number
- n
- integerMust be the true denominator of THIS percentage, not the study total.
- decimalsoptional
- integer
- data_typeoptional
- stringDefaults to unknown deliberately, so an omission refuses rather than assumes.
Output
Whether the percentage is achievable at that sample size. Means, concentrations and ratios are refused — they are not constrained to multiples of 100/n. A truncation-only pass caps at suspicious.
Reference
Brown, N. J. L., & Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. Social Psychological and Personality Science 8(4), 363–369. 2017.
GRIM sweep
grim_sweepWhen it applies
Several percentages that are supposed to share one denominator — a subgroup breakdown, or a row of a demographics table.
How it works
GRIM normally needs the sample size. Here the sample size is what is missing or disputed, so the logic runs backwards: for each reported value, enumerate every candidate n at which that value is achievable, then intersect those sets across all the values. What survives is the set of sample sizes that could have produced the whole row at once. An empty intersection is repeated under the truncation convention before any verdict is formed.
Inputs
- values
- list of numbers
- n_max
- integer
- n_max_sourceoptional
- stated / logically_necessary / assumedLoad-bearing, and it gates the verdict together with n_min: only a stated n_max plus an n_min reaching the floor licenses impossible.
- n_minoptional
- integer
- decimalsoptional
- integer
Output
The viable sample sizes per value and the joint intersection. A single surviving n is the recovered denominator; several mean the row does not pin it down. An empty intersection is NOT automatically a finding, and three outcomes are distinguished. If the values do share an n under TRUNCATION — which SPSS and several journal styles use — the verdict is consistent and it is recorded as a reporting convention, because a truncating reporter is not an impossibility. If no convention rescues them, impossible requires that the searched range provably bracket the true n: n_max must be the paper's own stated total and n_min must reach the floor. Otherwise the result is indeterminate, and says so — an empty intersection inside a range WE chose is a fact about the search, not about the paper.
Reference
Brown, N. J. L., & Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. Social Psychological and Personality Science 8(4), 363–369. 2017.
GRIMMER
grimmer_checkWhen it applies
A standard deviation reported beside a mean and sample size for integer data.
How it works
Standard deviations are granular for the same reason means are. GRIMMER enumerates the achievable pairs of sum and sum-of-squares behind the reported mean, subject to the parity constraint that the two must agree modulo 2, and asks whether any of them produces the reported SD at its printed precision.
Inputs
- mean
- number
- sd
- number
- n
- integer
- decimals / decimals_sdoptional
- integersPrinted precision of each. A wrong assumed precision accuses.
- scale_min / scale_maxoptional
- integers
- discreteness_sourceoptional
- stated / logically_necessary / assumed
Output
Consistent, or the constraint that failed. Only integer and parity failures license impossible; range-dependent failures on assumed bounds cap at suspicious, as do granularity modes on assumed discreteness. Both the sample and population SD conventions are tried, and a population-only match caps at suspicious.
Reference
Anaya, J. The GRIMMER test: A method for testing the validity of reported measures of variability. PeerJ Preprints 4, e2400v1. 2016.
DEBIT
debit_checkWhen it applies
A binary (0/1) variable reported as a mean and an SD — very common in Table 1 of a clinical paper.
How it works
For binary data the SD is not free: the mean fixes it exactly. DEBIT checks the reported SD against the value the mean and n force, testing both the sample (n−1) and population conventions before saying anything.
Inputs
- mean
- number (a proportion)
- sd
- number
- n
- integer
- decimals_mean
- integerPrinted decimals. A defaulted 2 dp falsely accused 73.6% of 1 dp rows in testing.
- decimals_sd
- integer
- data_typeoptional
- binary / proportionAnything else is refused rather than judged — a tool pointed at the wrong data must refuse, not accuse.
Output
Consistent, or an inconsistency with the SD the data would have to have. A row consistent under the population-SD convention is suspicious, never impossible — that convention is legitimate, and treating it as fraud produced 66 false accusations in a 450-case validation grid.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
DEBIT (whole table)
debit_check_batchWhen it applies
An entire demographics table of binary variables, screened in one pass.
How it works
Runs the same arithmetic over every row, then judges the table rather than the row. Escalation requires a pattern of impossible rows; a lone hit stays suspicious.
Inputs
- rows
- table of rows (mean, SD, n, printed decimals)
- data_typeoptional
- binary / proportionRequired in practice — omitting it refuses the whole batch.
- decimals_mean / decimals_sdoptional
- integersSupply per row; the batch default of 2 dp falsely accuses 1 dp proportions.
Output
Every non-consistent row plus an aggregate verdict, and always the denominator — three suspicious rows out of 140 checks and three out of four are very different papers. A batch in which no row could be evaluated is indeterminate, never consistent.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
2×2 reconstruction
reconstruct_2x2When it applies
A contingency result reported only as two column percentages and a total N, with the underlying counts withheld.
How it works
Enumerates every column split of the total and keeps the 2×2 tables whose cell counts round to both printed percentages. If a reported χ² is supplied, the candidate set is filtered by it as well.
Inputs
- pct_col1 / pct_col2
- numbers
- n_total
- integer
- decimalsoptional
- integer
- reported_chi2optional
- number
Output
The set of reconstructable tables, ordered by group-balance plausibility, or a proof that none exists. Impossibility is decided over every column split — the balance ratio only orders the survivors. A non-matching χ² yields indeterminate, never an accusation.
GRIM-U
grimu_checkWhen it applies
A p-value attributed to a Mann–Whitney U or Wilcoxon rank-sum test, with both group sizes known. Nothing else — t, F and χ² p-values are refused.
How it works
The U statistic takes only (half-)integer values between 0 and n₁·n₂, so at given group sizes only a finite set of p-values is achievable. The modelled set spans the normal approximation with and without continuity correction, plus the exact permutation p.
Inputs
- n1 / n2
- integers
- reported_p
- text, e.g. “0.043” or “<0.001”Passed as text so a threshold keeps its “<” — stripping it has turned an indeterminate result into a headline accusation.
- test_typeoptional
- mann_whitney / wilcoxonAny other test is refused, not scored.
- one_sidedoptional
- true / false
Output
Impossible below the U = 0 floor under every convention; suspicious inside a granularity gap; otherwise consistent. Known limitation, stated in the result: the tie-corrected variance that R and scipy use on tied data is not in the modelled set, so an impossible verdict on heavily tied data is unproven until checked again by hand.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
GRIM-U coexistence
grimu_coexistenceWhen it applies
Two or more nearly identical rank-sum p-values reported at the same group sizes — say 0.171 and 0.172.
How it works
Because the achievable p-values are a finite, unevenly spaced set, two p-values differing by less than the local spacing cannot both be real. This maps each reported value to the achievable U's and asks whether they can coexist.
Inputs
- n1 / n2
- integers
- p_values
- list of numbers
- one_sidedoptional
- true / falseForcing the two-sided set on one-tailed reports manufactures accusations.
Output
Whether the reported values can all arise at these group sizes. Carries the same tie-correction limitation as GRIM-U.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
SPRITE
sprite_reconstructWhen it applies
A mean and SD on a bounded integer scale, when you want to see what the underlying data could have looked like — or prove it could not have existed.
How it works
Searches for integer samples on the stated scale whose mean and SD round to the reported values, matching the SD at its reported precision rather than on exact variance. Alongside the search runs an analytic integer-and-parity screen that can prove no sample exists.
Inputs
- mean
- number
- sd
- number
- n
- integer
- scale_min / scale_maxoptional
- integers
- decimals_mean / decimals_sdoptional
- integers
- discreteness_sourceoptional
- stated / logically_necessary / assumed
Output
One of four outcomes, and only one of them is evidence: solutions found, no solution proven (from the analytic constraint), GRIM failure, or search exhausted. An exhausted search is a failed search, never an anomaly — conflating the two is how a reconstruction tool becomes an accusation generator. Recovered distributions with implausible shapes are returned for a human to judge.
Reference
Heathers, J. A., Anaya, J., van der Zee, T., & Brown, N. J. L. Recovering data from summary statistics: Sample Parameter Reconstruction via Iterative TEchniques (SPRITE). PeerJ Preprints 6, e26968v1. 2018.
Test-statistic recomputation10 tools
A test statistic, its degrees of freedom and its p-value are redundant with one another and with the group summaries they came from. Any one can be recomputed from the others, and the recomputed value has to agree with what was printed. These tools do that arithmetic — always against the interval the printed rounding allows, never against a point value, because matching a rounded input to an exact expectation is how a recomputation tool starts accusing honest papers.
statcheck
statcheckWhen it applies
Any paper reporting results in APA format — a free, broad first pass over the whole text.
How it works
Extracts APA-formatted test statistics from the prose and recomputes each p-value from its statistic and degrees of freedom, flagging inconsistencies and, separately, decision errors where the recomputed p falls on the other side of the significance threshold.
Inputs
- text
- the paper's textA PDF path is accepted as a fallback route.
Output
Every extracted result with its reported and recomputed p, marked consistent, inconsistent or a decision error. Each p is recomputed assuming a two-tailed unadjusted test, and only APA grammar is visible to it — everything else is invisible, so a clean run is not a clean paper. Reporting inconsistencies are near-universal in honest work and must never support a misconduct verdict alone.
Reference
Nuijten, M. B., Hartgerink, C. H. J., van Assen, M. A. L. M., Epskamp, S., & Wicherts, J. M. The prevalence of statistical reporting errors in psychology (1985–2013). Behavior Research Methods 48(4), 1205–1226. 2016.
Independent-samples t-test
recalc_independent_tWhen it applies
Two groups' means, SDs and sizes reported alongside a t or a p.
How it works
Recomputes t and p from the six group statistics and compares them against the interval the printed precision allows. Both Welch's and Student's pooled readings are accepted, and both tail conventions, unless the paper states which it ran.
Inputs
- mean1, sd1, n1
- numbers and an integer
- mean2, sd2, n2
- numbers and an integer
- reported_t / reported_poptional
- numbers
- use_welchoptional
- true / falseOmit when the paper does not say; both standard tests are then accepted.
- p_tailoptional
- one / two / unknown
- decimals_mean / decimals_sd / decimals_poptional
- integers
Output
The recomputed t and p with the achievable interval, and a verdict on the reported values. Convicting a pooled t against a Welch recomputation is a false accusation, which is why neither reading is assumed.
Within-subjects t-test
recalc_within_subjects_tWhen it applies
A paired (pre/post) design reporting both timepoints' means and SDs plus a t or p, where the correlation between them is not stated.
How it works
A paired t implies a specific pre/post correlation, which can be recovered algebraically from the summary statistics. A recovered correlation outside [−1, 1] describes data that cannot exist. Because a thresholded p only lower-bounds t, the recovered r is then a lower bound too.
Inputs
- mean_pre, sd_pre, mean_post, sd_post
- numbers
- n
- integer
- reported_p
- text, e.g. “0.03” or “<.05”Passed as text so a threshold keeps its “<”. Stripping it once turned an indeterminate interval into a headline finding against a real paper.
- p_is_thresholdoptional
- true / false
- p_tailoptional
- one / two
Output
The recovered correlation and the whole interval attainable across the inputs' rounding box. Flagged only when that entire interval lies outside [−1, 1]; an exact-p impossibility caps at highly suspicious, because a reporting convention rather than fabrication is the usual cause.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
One-way ANOVA
recalc_anovaWhen it applies
A one-way ANOVA reported with its group means, SDs and sizes.
How it works
Rebuilds the between- and within-group mean squares from the group summaries, recomputes F and p, and compares them against the interval the printed rounding allows — rounding propagates multiplicatively into F, so a point comparison would be wrong.
Inputs
- means / sds / ns
- lists of numbers
- reported_f / reported_poptional
- numbers
- decimals_mean / decimals_sd / decimals_poptional
- integers
Output
Recomputed F and p with their achievable intervals. Assumes a classical fixed-effects one-way ANOVA rather than Welch's, with homogeneous variances and independent observations — all of which are stated in the result.
Chi-squared
recalc_chi_squaredWhen it applies
A contingency table whose counts are printed alongside a χ² or a p.
How it works
Recomputes χ² from the observed counts under both the Pearson and Yates conventions, and also computes Fisher's exact p for 2×2 tables. All three are tried before anything is flagged.
Inputs
- observed
- table of counts (rows × columns)
- reported_chi2 / reported_poptional
- numbers
- use_yatesoptional
- true / false
Output
Each convention's statistic and p against the reported values. The likelihood-ratio (G-test) convention is deliberately not tried and the result says so — a paper reporting G can mismatch entirely innocently.
F → p
recalc_f_to_pWhen it applies
Any reported F with both degrees of freedom and a p — usable on ANOVAs of any shape, not just one-way.
How it works
The printed F is rounded, so it asserts a range of p rather than a value. This computes that range from the F's printed precision and its degrees of freedom, and checks the reported p against it.
Inputs
- f_value
- number
- df1 / df2
- integers
- reported_poptional
- number
- decimals_f / decimals_poptional
- integers
Output
The implied p-range and a verdict. Directional: only a p below the achievable range is flagged, since an under-claimed p is conservative reporting rather than an error.
Any statistic → its p-value
recalc_stat_to_pWhen it applies
A single printed t, χ², z, r or F with a p beside it — the general-purpose version of the checks above.
How it works
Recomputes the p from the named distribution and degrees of freedom. Because the statistic is rounded it implies a p-interval, and the reported p is judged against that interval rather than a point.
Inputs
- statistic
- number
- kind
- t / chi2 / z / r / F
- df / df2 / noptional
- integersWhichever the named test needs; a missing one is named in the refusal.
- reported_poptional
- number
- tailsoptional
- 1 or 2Ignored for χ² and F.
Output
The implied p-interval and a directional verdict — only an over-claim is flagged. Capped at suspicious, tighter than its class allows: a one-tailed test, a different df convention or an unshown correction would each explain a disagreement. A missing df or n is reported as an incompletely reported test, which is a property of the paper.
Regression coefficient
check_regressionWhen it applies
A regression table printing a coefficient, its standard error, and a t or p.
How it works
Checks that the reported t equals B divided by SE, judged against the interval reachable from the printed precision of B and SE, and that the p is that t's tail probability at the given degrees of freedom. Both sign conventions (signed t and |t|) and both tail conventions are accepted first.
Inputs
- b
- number
- se
- number
- reported_t / reported_poptional
- numbers
- dfoptional
- integerWithout it only the p check is skipped; t = B/SE still runs.
- p_is_thresholdoptional
- true / falseTrue when the paper printed a bound such as “P < .001” — a float cannot carry that difference.
Output
The implied t, its achievable interval, and verdicts on the reported t and p separately.
Regression table (whole)
check_regression_batchWhen it applies
Every coefficient in a regression table, in one call.
How it works
Runs the coefficient check per row. Where the degrees of freedom were not reported it re-runs at df of 10, 30, 100 and 100,000; a verdict that changes across that range is returned as indeterminate rather than resolved by a guess.
Inputs
- rows
- table of rows (b, se, reported t/p, df, printed decimals, label)
Output
Per-row results plus an aggregate set by the worst row. A batch in which no row could be evaluated is indeterminate, never consistent — and the denominator is always reported alongside the hits.
STALT — hidden p-values
stalt_checkWhen it applies
A paper reporting only “p < 0.05” where the group statistics permit an exact recalculation.
How it works
Compares the recomputed exact p against the threshold the paper printed. The threshold is read as an upper-bound claim, so any p at or below it is consistent; what the check surfaces is a p many orders of magnitude below the stated bound — information the paper had and did not report.
Inputs
- calculated_p
- numberMust come from the same test the printed threshold describes.
- reported_thresholdoptional
- text, e.g. “<0.05”
Output
The size of the overshoot and a severity that scales with it. A factor-of-two overshoot is tolerated as a one- versus two-tailed artefact rather than flagged.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
Table arithmetic & internal consistency10 tools
Papers state the same quantity more than once, in different forms, and the forms have to agree. A count and its percentage share a denominator; subgroups sum to their total; an estimate sits at the centre of its own confidence interval; a standard deviation cannot exceed what the variable's range permits. None of this needs raw data — only that the numbers printed in one table be consistent with the numbers printed beside them.
Whole-table check
check_tableWhen it applies
Any table — this is the entry point, and it replaces calling the per-cell tools one at a time.
How it works
A router. Given typed cells, the variables you declared and a binding from row labels to variables, it determines which checks apply from cell kind crossed with variable family and runs all of them. Which checks apply is derived, not chosen. Declaring a variable also resolves cells that syntax alone must refuse: “48.90 (14.46)” under a bare label is ambiguous until the variable is declared continuous.
Inputs
- cells
- typed table cells
- variablesoptional
- list of variable declarationsDeclare from the paper, never from a guess — a fabricated bound inverts a check rather than weakening it.
- bindingsoptional
- map of row label → variable
- provenanceoptional
- printed_table / measured_raster / ocr_table / …Figure-read and OCR'd values make the exact-arithmetic checks refuse automatically.
Output
Every non-consistent row in full, a count of the consistent ones, and a machine-generated list giving the reason each check was refused — so coverage does not depend on anyone remembering. It also returns a declaration template naming exactly which variables are blocking checks it could otherwise run. The impossible ceiling is the router's, not every row's: read each row's own tool to know what its verdict rests on.
Count against percentage
check_count_percentWhen it applies
Any “n (%)” cell — which is most of Table 1 in a clinical paper. Use it whenever the count is printed; it is strictly stronger than GRIM on a percentage.
How it works
Divides the printed count by the printed denominator and checks the result against the printed percentage, at the rounding interval the percentage's own precision asserts.
Inputs
- count
- integer
- percentage
- number
- n
- integer
- decimalsoptional
- integerWithout it the check returns indeterminate rather than assuming a precision.
Output
A verdict plus the implied denominator — and across a table the implied denominators are the real evidence, because a quantity that can only have one value should not imply several. Capped at suspicious, tighter than its class allows: row, column and subgroup bases legitimately coexist in one table, and no arithmetic disagreement rules that out. A count exceeding its denominator is reported as a wrong pairing on our side, not a defect in the paper.
Ratio check
check_ratioWhen it applies
A printed ratio shown with its own numerator and denominator — cost-effectiveness ratios, rates, “X per Y”.
How it works
Computes the quotient at the corners of the rounding box implied by all three printed precisions, and asks whether the printed ratio falls inside it.
Inputs
- numerator
- number
- denominator
- number
- reported_ratio
- number
- decimals_numerator / decimals_denominator / decimals_ratiooptional
- integersAll three are needed; without them the check refuses, because a guessed precision either clears every ratio or accuses on rounding alone.
Output
A verdict plus the implied numerator and denominator. Capped at suspicious despite testing exact arithmetic: papers routinely divide unrounded intermediates they never show, or a discounted or subgroup quantity. Across a table, several rows implying different values for a quantity that can only have one is the finding — not any single row. A denominator whose rounding interval spans zero makes the quotient unbounded, and the check refuses rather than compares.
Summation check
check_summationWhen it applies
Values that should add to a stated total — subgroup counts to N, percentages to 100, cost components to a total cost.
How it works
Sums the addends and compares against the stated total within a window derived from the printed decimals of every addend and of the total, so rounding is accounted for rather than assumed away.
Inputs
- values
- list of numbers
- expected_sum
- number
- decimalsoptional
- list of integersPrinted decimals per addend. With neither this nor an explicit tolerance the check refuses rather than guessing.
- toleranceoptional
- numberCount columns that must hold exactly take 0; percentage columns legitimately misbalance by rounding.
Output
The computed sum, the permitted window and the discrepancy. The values are assumed to be exhaustive, non-overlapping addends — an omitted row or an overlapping category mimics a mismatch, and the result says so.
Summation (whole table)
check_summation_batchWhen it applies
Every summation in a table, in one call.
How it works
Runs the same check per row, with per-row printed decimals, and sets the aggregate verdict from the worst row.
Inputs
- rows
- table of rows (values, expected total, printed decimals)
- toleranceoptional
- numberA batch-wide override.
Output
Per-row results and an aggregate. A row with no printed precision is refused rather than guessed, and a batch in which nothing could be evaluated is indeterminate, never consistent.
Value against its logical range
check_statistic_boundsWhen it applies
A value that appears to fall outside what its measure permits by definition — a correlation above 1, a p-value above 1, a percentage above 100.
How it works
Only bounds that follow from a measure's definition are known to it — a correlation is bounded by Cauchy–Schwarz, a p-value by the definition of a probability. A merely expected range is an assumption, so unknown measures are refused rather than judged against our expectations. Crucially, it treats a violation as evidence about our extraction before evidence about the paper.
Inputs
- value
- number
- measure
- correlation / p_value / percentage / …Anything without a definitional bound is refused, not guessed at.
- extraction_verifiedoptional
- true / falseSet only after re-reading the value in the paper and confirming both the digits and the column. This is what lifts the refusal to impossible.
Output
By default, indeterminate with the failure mode “extraction suspect”, naming the decimal-point slip or wrong-column read that would explain it — because gross violations are uncommon even in fraudulent papers, since fabricators produce plausible numbers. Verification is the price of the accusation. Even when verified, the tool names the innocent readings, because a confirmed impossible value still does not establish who produced it.
SD range check
check_sd_rangeWhen it applies
A standard deviation for a variable whose attainable range is known.
How it works
The largest SD a bounded sample can have is fixed by its range — (range/2)·√(n/(n−1)) for the sample SD. A reported SD above that ceiling describes data that cannot exist. With n unknown the ceiling is taken at its weakest over all n; supplying n tightens it.
Inputs
- sd
- number
- min_val / max_val
- numbersMust be the variable's true attainable bounds.
- noptional
- integer
- bound_sourceoptional
- stated / logically_necessary / assumedLoad-bearing: only a range the paper states, or one that is logically necessary, licenses impossible. An assumed range caps at highly suspicious.
Output
The ceiling, the reported SD and the verdict. A negative SD is bound-free and impossible whatever else is unknown.
SD-or-SE adjudication
check_sd_or_seWhen it applies
A dispersion value where it is unclear whether the paper reported a standard deviation or a standard error.
How it works
SD and SE are linked by SE = SD/√n, so both readings can be tested against a plausibility ceiling implied by the variable's range and the reading that fits adjudicated. Under the SE reading the rounding interval stretches by √n, which is accounted for.
Inputs
- reported_value
- number
- n
- integerMust be the cell's true denominator.
- reported_asoptional
- sd / se / unknown
- variable_range_min / variable_range_maxoptional
- numbers
- bound_sourceoptional
- stated / logically_necessary / assumed
Output
Which reading the value must be, or that neither fits. A “neither fits” verdict rests entirely on the plausibility ceiling, so on an assumed bound it caps at suspicious. The result carries the base rate: mislabelling SE as SD is the norm rather than an outlier — Olsen found 35 of 88 studies in one journal doing it — and it is usually a statistical error, not a sign of misconduct.
Reference
Olsen, C. H. Review of the Use of Statistics in Infection andImmunity. Infection and Immunity 71(12), 6689–6692. 2003.
Estimate against its own interval
check_estimate_ciWhen it applies
A point estimate printed with a confidence interval, and often a p-value too — odds ratios, hazard ratios, mean differences.
How it works
Checks ordering, containment and the midpoint on the correct scale — linear for a difference, logarithmic for a ratio measure — and recovers the p implied by the interval to compare against the reported one. Alternative confidence levels are tried before any “wrong p” is reported.
Inputs
- estimate
- number
- ci_low / ci_high
- numbers
- measureoptional
- OR / RR / HR / IRR / differenceDecides whether symmetry is checked on the linear or the log scale.
- reported_poptional
- number
- reported_p_operatoroptional
- =, <, ≤, >, ≥Pass “<” for “P < .001” — a threshold is an interval claim, not a value.
Output
Each sub-check separately with its verdict. Exact, profile and bootstrap intervals are legitimately asymmetric and can fail the midpoint check innocently, which the result states. A p that cannot be placed against alpha at its own printed precision is left unclassified rather than judged.
Continuous → categorical plausibility
check_proportion_from_normalWhen it applies
A paper reporting both a continuous variable's mean and SD and the proportion of participants above or below some threshold on it.
How it works
Given the mean, SD and n, the share of the sample past a threshold is constrained. The headline test assumes the raw variable is normal and uses a binomial test on the count at the threshold; a second, distribution-free test uses Cantelli's inequality, which holds for every distribution.
Inputs
- reported_proportion
- number
- mean / sd
- numbers
- n
- integer
- threshold
- number
- directionoptional
- above / below
Output
Both tests separately. The normal-assumption result caps at suspicious, because skew explains a mismatch innocently; only a violation of the distribution-free Cantelli bound reaches highly suspicious. The proportion and the mean/SD/n must describe the same sample.
Similarity & duplication5 tools
Real data is noisy in ways people reliably fail to imitate, and it is never noisy twice in the same way. Randomised groups differ at baseline by chance; sample standard deviations scatter by a knowable amount; two independent samples do not produce the same mean and SD. Fabricators add noise to the means, because everyone knows means vary, and forget that the dispersion measures need their own — and when a table or a spreadsheet is filled by copying, the duplicate values survive in the published numbers. The first three tools here test for data that is too tidy; the last two are copy-paste detectors, working on printed summary statistics and on participant-level raw data respectively. None of them can ever return “impossible”: no arrangement of numbers is forbidden by arithmetic, only improbable.
Carlisle–Stouffer–Fisher test
csf_testWhen it applies
A randomised trial's Table 1, where baseline characteristics are compared between arms.
How it works
Under simple randomisation, baseline p-values are uniform and independent. Combining them by Stouffer's or Fisher's method gives a one-sided test for EXCESS similarity. A very small result means the arms are more alike than randomisation can plausibly produce — the signature of a trial whose allocation never happened.
Inputs
- p_values
- list of numbersAt least three usable values are needed.
- designoptional
- simple / cluster / stratified / crossover / …Cluster, multi-site, stepped-wedge, crossover, matched, stratified and minimised designs force baseline balance by construction and are refused outright. An unstated design caps the verdict.
- correlationoptional
- numberBaseline rows are correlated; with none supplied an equicorrelated null is read at a conservative ρ = 0.5.
Output
A one-sided similarity p-value with a sensitivity profile across correlation assumptions. High baseline p-values are the red flag here, never reassurance — a Table 1 where every p is 0.93, 0.99, 0.99 is the thing this test exists to catch.
References
Carlisle, J. B. Data fabrication and other reasons for non‐random sampling in 5087 randomised, controlled trials in anaesthetic and general medical journals. Anaesthesia 72(8), 944–952. 2017.
Carlisle, J. B. False individual patient data and zombie randomised controlled trials submitted to Anaesthesia. Anaesthesia 76(4), 472–479. 2021.
Bayesian Table-1 dispersion
bayesian_table1_dispersionWhen it applies
The same Table 1 as the CSF test — complementary to it, not preferred over it, and handling mixed continuous and categorical rows.
How it works
Turns each baseline row into a p-value and then a z-score, and estimates a precision multiplier τ̂ = √(1/mean(z²)). Above 1 indicates under-dispersion — differences systematically smaller than chance allows; below 1, over-dispersion and possible randomisation failure. Rows are treated as equicorrelated, discounting a table of any size to about two effective degrees of freedom, so a large table is not thereby more accusable.
Inputs
- rows
- list of baseline rows — means/SDs/ns, or categorical counts
- designoptional
- stringSame refusals as the CSF test; an unstated design caps the verdict below highly suspicious.
- correlationoptional
- numberPass only if genuinely known or estimable.
Output
τ̂ with a confidence interval that inverts the exact chi-square null, plus a sensitivity profile. An elevated τ̂ whose interval does not exclude 1 is reported indeterminate, never consistent — too few effective degrees of freedom to convict OR to clear. Escalation additionally requires the CSF test on the same p-values to corroborate; CSF finding nothing vetoes it.
Reference
Bolland, M. J., Gamble, G. D., Avenell, A., Grey, A., & Lumley, T. Baseline P value distributions in randomized trials were uniform for continuous but not categorical variables. Journal of Clinical Epidemiology 112, 67–76. 2019.
Variance dispersion
check_variance_dispersionWhen it applies
Three or more groups whose standard deviations are reported — are the SDs too similar to one another?
How it works
Sample variances are themselves random: run the same experiment on several groups and the SDs scatter by a knowable amount. Standardising by the within-group mean square leaves only relative spread, whose null is bootstrapped from the chi-square distribution. Dispersion is measured as the range of the z-scores rather than their SD, which the source study found more robust to heterogeneity.
Inputs
- sds / ns
- lists of numbersAt least three groups.
- sd_texts
- list of the SDs as printedPrecision is read from the literal; rounding coarser than a quarter of the SD is refused, because coarse rounding drives honest SDs onto identical values.
- homogeneous_variancesoptional
- true / falseRequired — a violated assumption blinds the test rather than making it accusatory, so it gates the consistent verdict too.
- designoptional
- stringDesigns that constrain variances by construction are refused.
Output
The observed dispersion against its bootstrapped null, capped at suspicious — over-similar SDs are improbable, never impossible. Every result sets a reference-required flag: the validating study's near-perfect discrimination came from comparing genuine against fabricated sets, while both looked significant against the theoretical null in absolute terms. This is a relative verdict, and the result says so.
References
Hartgerink, C. H. J., Voelkel, J. G., Wicherts, J. M., & van Assen, M. A. L. M. Detection of data fabrication using statistical tools. PsyArXiv . 2019.
Simonsohn, U. Just Post It: The Lesson From Two Cases of Fabricated Data Detected by Statistics Alone. Psychological Science 24(10), 1875–1888. 2013.
Summary-statistic copy-paste
detect_duplicationWhen it applies
Across a paper's tables and studies, wherever summary statistics are supposed to describe independent samples — and within a single table, where a filled-down row or column leaves the same values behind.
How it works
Hashes every (mean, SD) cell at its PRINTED precision and looks for three signatures. Across blocks the paper calls independent samples, it counts identical pairs and prices them against a Poisson chance model whose value spans are inferred from the data itself rather than assumed. Within a single table, it looks for a whole row or column of cells repeated under a different label — the fill-down signature. And it reports pairs identical after a single digit transposition, as a weak pointer only.
Inputs
- blocks
- labelled blocks of cells, each with a sample identityBlocks are assumed independent unless they share a sample id — Tables 1, 2 and 3 of a cohort paper are usually subsets of each other.
- decimals_mean / decimals_sdoptional
- integersThe printed precision cells are hashed at. Cells may override per-cell.
- near_matchoptional
- true / falseAlso report pairs identical after a digit transposition. Always indeterminate — a pointer, never a finding.
Output
Each duplicate group with its kind, the matched values, and a multiplicity-corrected chance probability. Severity keys on the number of identical pairs — three escalates — and several guards pull it back down. Blocks sharing a sample id, or declared non-independent, drop two tiers. Round values (integers, halves, multiples of 0.05) collide far more often and drop one. A lone match is re-checked against its own chance probability and demoted if a coincidence was likely. Within a table, integer count cells with no SD never count toward severity at all: blocked allocation and zero attrition make arm sizes identical BY DESIGN. Capped at highly suspicious — duplication is improbable, never impossible.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
Raw-data copy-paste
detect_raw_duplicationWhen it applies
A participant-level supplement — a spreadsheet of rows, one per subject. The strongest evidence available, when the data exists.
How it works
Takes a file path and loads the grid itself — a raw sheet is thousands of values and no model can retype it — then looks for three signatures: duplicated participant rows within and across sheets, vertical runs of values reappearing in order, and improbably specific recurring numbers. Matching is exact after float canonicalisation; there is no tolerance anywhere, because a tolerance manufactures matches between genuinely different measurements. Evidence is weighted by the INFORMATION CONTENT of each number rather than by match count: 123.46 recurring means something, 100 recurring means nothing, and a value shared by twenty rows is a site-level covariate rather than twenty copy-pastes.
Inputs
- path
- spreadsheet file (.xlsx, .csv, .tsv, .docx)
- sheetoptional
- worksheet name
- exclude_columnsoptional
- list of column namesColumns shared by design — IDs, site codes, coordinates, plot-level covariates — must be excluded by the caller or they duplicate legitimately.
- data_contextoptional
- raw_participant_level / pooled_or_meta_analysis / unknownA pooled or meta-analytic file duplicates rows by construction; findings are then capped at indeterminate. That is a refusal, not a finding.
Output
Each signature with its own chance model — a Poisson null over same-precision bins for repeated values, an empirical collision null for matching runs — so a large sheet is not flagged merely for being large. Row-pair evidence has no chance model and the result says so plainly; it rests on information content alone and is the weaker limb. Capped at highly suspicious.
Reference
Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.
Design, power & inference5 tools
Not every problem in a paper is a wrong number. A stated power analysis can be recomputed against the sample actually recruited; a family of tests has an arithmetic false-positive rate whatever the paper says about it; a “no difference” conclusion can be checked against the effect its own confidence interval still permits. These tools audit the inferences rather than the arithmetic — and several of them are calculators that render no verdict at all.
Power-analysis verification
verify_power_analysisWhen it applies
A paper stating an a-priori power analysis — the effect size, alpha and target power it planned around.
How it works
Recomputes the required sample size using exact noncentral distributions and compares it against the N actually recruited, with a small tolerance absorbing documented differences between G*Power, pwr and PASS.
Inputs
- test_type
- string
- effect_size / alpha / target_poweroptional
- numbers
- n_reportedoptional
- integer
- tailsoptional
- 1 or 2As stated by the paper. Left unset, both readings are tried and no accusation is made when the one-tailed reading reconciles.
- allocation_ratiooptional
- numberAnything other than 1.0 is refused — the recomputation is 1:1, and grading a k:1 design against it can accuse a correct calculation.
Output
The required N against the recruited N. Indeterminate unless every input was stated — which is the modal honest outcome and is itself worth reporting: roughly 80% of published power analyses cannot be checked at all, because the inputs are not there. Over-recruiting is treated as correct practice, never a discrepancy. It cannot detect an analysis that reproduces but powers the wrong hypothesis.
Reference
Bakker, M., Veldkamp, C. L. S., van Assen, M. A. L. M., Crompvoets, E. A. V., Ong, H. H., Nosek, B. A., Soderberg, C. K., Mellor, D., & Wicherts, J. M. Ensuring the quality and specificity of preregistrations. PLOS Biology 18(12), e3000937. 2020.
Required sample size
required_nWhen it applies
To ask what sample a design would have needed — as context for a study, not as a charge against it.
How it works
A pure calculator: the smallest n reaching a target power for a named test at a given effect size and alpha, via exact noncentral distributions. Supplying the sample actually used adds a design-sensitivity comparison — the effect the study could have detected.
Inputs
- test_type
- string
- effect_size
- number
- alpha / poweroptional
- numbers
- actual_noptional
- integerAdds the design-sensitivity block; without it the answer is a bare sample-size calculation.
Output
The required n, and optionally the minimum detectable effect. It renders no verdict about any paper: a small study is a fact about sensitivity, not a defect.
Multiple-comparisons arithmetic
multiplicity_reportWhen it applies
A paper running many tests and treating each at α = 0.05 — or claiming a correction that can be checked.
How it works
For a declared family of tests, computes the family-wise error rate and the thresholds four standard corrections would impose: Bonferroni, Holm, Benjamini–Hochberg and Benjamini–Yekutieli. It then counts how many reported claims survive each.
Inputs
- tests
- list of tests (p-value, whether a correction was applied)
- family_label
- textDescribe the family you chose. Family definition is a judgement the tool cannot make, so there is no default — an omitted label is refused, not filled in.
- alphaoptional
- number
Output
The family-wise error rate, the four thresholds, and floor and ceiling counts of surviving claims. Everything is conditional on the declared family — the same p-values give different answers under different families — and the survivor count is a floor, because undisclosed analytic flexibility means the true comparison count can exceed the reported one. Deterministic arithmetic; it accuses no one.
References
Benjamini, Y., & Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B 57(1), 289–300. 1995.
Benjamini, Y., & Yekutieli, D. The control of the false discovery rate in multiple testing under dependency. The Annals of Statistics 29(4), 1165–1188. 2001.
Absence of evidence
absence_of_evidenceWhen it applies
A paper concluding “no effect” or “no difference” from a non-significant result.
How it works
“Not significant” is not “no effect”. This reports the largest effect the paper's own confidence interval still permits, and judges the CLAIM made about the interval rather than the data. If no interval is printed, one is derived from the coefficient and standard error and marked as derived so it never passes as the paper's.
Inputs
- ci_low / ci_highoptional
- numbersOr pass b and se instead and an interval is derived.
- measureoptional
- OR / RR / HR / IRR / differenceRequired in practice: the null is 1 for a ratio measure and 0 for a difference, and without it the tool declines. Omitting it once inverted a verdict — the interval [0.85, 1.42] was reported as excluding the null.
- sesoioptional
- numberSmallest effect size of interest. Without one, equivalence is described but not judged — that cannot be claimed against an unstated threshold.
- dfoptional
- integerWithout it the normal distribution is used and the result says so.
Output
The interval, the largest effect it still admits, and whether the stated conclusion survives it. Declines rather than guesses when the null cannot be located or no threshold was supplied.
Effect-size plausibility
check_effect_plausibilityWhen it applies
Any two-group comparison — and notably the one check that still works on figure-derived data, so a paper reporting its outcome only as a bar chart becomes checkable here.
How it works
Computes Cohen's d and Hedges' g with a confidence interval from the six group statistics. Passing the paper's own reported d alongside them checks internal consistency instead: does the printed effect size match the one its own means, SDs and ns imply? That needs no comparison class and is what catches a mistyped effect size.
Inputs
- mean1, sd1, n1, mean2, sd2, n2optional
- numbers and integersOr pass d directly.
- reported_doptional
- numberTriggers the internal-consistency check.
- fieldoptional
- stringUsually omitted. Adds field benchmarks — but only social psychology's percentiles trace to a published table; the rest are uncited estimates and say so.
- provenanceoptional
- printed_text / figure_pixels
Output
The effect size with its interval — reporting the value is the deliverable. Capped at suspicious: an effect can be improbable but never arithmetically impossible, and an unusually large one is informative, not evidence of misconduct. Before treating a large effect as a finding, three ordinary causes are named in the result: SEM read as SD (which inflates d by √n), a waitlist rather than active comparator, and a units mismatch. Field benchmarks are distributions of published effects, themselves inflated by publication bias.
Reference
Lovakov, A., & Agadullina, E. R. Empirically derived guidelines for effect size interpretation in social psychology. European Journal of Social Psychology 51(3), 485–504. 2021.
Hand-calculation signatures & heuristics4 tools
Some anomalies are not impossible numbers but improbable patterns: a test statistic that exactly equals the value you would get by hand from the rounded table, sample sizes too round to be real recruitment, or a comparison reported so coarsely that no one can tell what was computed. These are triage signals. They escalate alongside stronger findings and are capped so they cannot convict on their own.
RIVETS (t-test)
rivets_independent_tWhen it applies
A t-test reported beside the group summaries it supposedly came from — the precision-aware replacement for a plain recomputation.
How it works
Real data has more precision than the table prints, so a t computed from actual data almost never lands exactly on the value you get from the rounded summaries. RIVETS maps the t and p reachable across the inputs' rounding intervals by Monte Carlo, and measures how often an exact match occurs. A point match plus a rare hit rate is the signature of a statistic computed from the table rather than from data.
Inputs
- mean1, sd1, n1, mean2, sd2, n2
- numbers and integers
- reported_t / reported_poptional
- numbers
- decimals_mean / decimals_sd / decimals_t / decimals_poptional
- integersMust be supplied; the defaults assume more precision than most papers report.
- p_is_thresholdoptional
- true / falsePass a bound p as text, never as a float, or it reads as an unreachable exact value.
Output
The achievable t and p intervals, the exact-match hit rate, and a verdict. Both Welch and Student readings and both tail conventions are accepted unless the paper states them.
Reference
Brown, N. J. L., & Heathers, J. Rounded Input Variables, Exact Test Statistics (RIVETS). PsyArXiv . 2019.
RIVETS (ANOVA)
rivets_anovaWhen it applies
The same detector applied to a reported F. Preferred over the plain ANOVA recomputation, which is precision-blind.
How it works
Monte Carlo over the truncation interval of every printed input, producing the reachable F and p ranges plus the exact-match hit rates.
Inputs
- means / sds / ns
- lists of numbers
- reported_f / reported_poptional
- numbers
- decimals_mean / decimals_sd / decimals_f / decimals_poptional
- integers
- p_is_thresholdoptional
- true / false
Output
Reachable ranges, hit rates and a verdict. A reported p outside the achievable range is flagged at the same severity as an out-of-range F. This tool owns the accusation-grade reported-p check for ANOVAs.
Reference
Brown, N. J. L., & Heathers, J. Rounded Input Variables, Exact Test Statistics (RIVETS). PsyArXiv . 2019.
Round-N flag
flag_round_nWhen it applies
A study reporting several distinct sample sizes that are all suspiciously round.
How it works
A base-rate heuristic. Fabricated papers over-represent round Ns — but so do honest recruitment targets, which is why it needs at least three reported sizes with at least two distinct before it says anything at all.
Inputs
- ns
- list of integers
- moderate_modulus / strong_modulusoptional
- integers
Output
A triage flag capped at suspicious — a weak signal that escalates only alongside stronger findings, and never a finding on its own.
Uninformative-statistic flag
flag_uninformative_statWhen it applies
A comparison reported so coarsely that the printed precision cannot pin down the test outcome.
How it works
Models the comparison as a one-way ANOVA over the group summaries and computes the whole p-range the printed precision permits — by Monte Carlo plus an exact corner search, because the extreme of F lives on a vertex that interior sampling reaches with probability zero. It then asks whether that range straddles the significance threshold.
Inputs
- means / sds / ns
- lists of numbers
- reported_poptional
- number
- decimals_mean / decimals_sd / decimals_poptional
- integers
- significance_thresholdoptional
- number
Output
The analytic p-range, with flags for a range straddling alpha, an implausibly wide range, or a reported p outside it. Capped at suspicious with the one-way-ANOVA assumption disclosed; the RIVETS ANOVA check owns anything accusation-grade.
Cross-checks against R2 tools
Two tools exist only to check our own arithmetic against the canonical R implementations. They render no verdict about any paper: a disagreement between engines means OUR tooling is wrong, and the finding is withheld rather than reported. An engine that cannot run is never a disagreement — the wrappers distinguish “package missing”, “R absent” and “R errored”, because a missing package once silently capped every GRIM finding in a whole corpus.
GRIM ground truth (R scrutiny)
r_grim_ground_truthWhen it applies
Before publishing any GRIM result at impossible — the prompts require this cross-check first.
How it works
Shells out to R and runs the same test through the scrutiny package, then compares verdicts.
Inputs
- mean
- number
- n
- integer
- itemsoptional
- integer
- decimalsoptional
- integer
Output
R's verdict beside ours. A disagreement is a bug on our side and is never evidence about the paper. If R or the package is unavailable the result says which — it does not report a disagreement.
SPRITE ground truth (R rsprite2)
r_sprite_ground_truthWhen it applies
To confirm a SPRITE no-solution result before it is relied on.
How it works
Runs the reconstruction through R's rsprite2 and reports whether it proves infeasibility. Note that rsprite2 infers decimal precision from the literal it is passed, so the precision must be pinned explicitly or the two engines are silently answering different questions.
Inputs
- mean / sd
- numbers
- n
- integer
- scale_min / scale_maxoptional
- integers
- decimalsoptional
- integerPin this explicitly — rsprite2 reads “4.0” as zero decimals.
- seedoptional
- integerVary it to test whether a no-solution result is robust.
Output
rsprite2's samples, or its proof that none exists. A disagreement with our implementation is a bug in our tooling and the finding must not be reported.
Reference
Heathers, J. A., Anaya, J., van der Zee, T., & Brown, N. J. L. Recovering data from summary statistics: Sample Parameter Reconstruction via Iterative TEchniques (SPRITE). PeerJ Preprints 6, e26968v1. 2018.
The toolkit as a whole follows James Heathers’ An Introduction to Forensic Metascience. Individual methods are credited on their own cards. See also the findings the toolkit has reported.