← return to the Forensic Metascience AI toolkit

The toolkit

This page documents 46 tools in the Forensic Metascience AI toolkit. Every one is a deterministic function, not a language model: it takes numbers the paper already printed, does arithmetic, and returns a structured result with a severity. The tools agent chooses which tools to point at a paper and how to read what comes back — it never decides what a check concludes.

None of these need the raw data. That is the point of the field: a paper's own summary statistics constrain each other tightly enough that fabrication and error leave traces in the numbers as published.

The tools agent’s registry holds 58 tools. The ones not listed here are left out on purpose, and for three different reasons. Ten read PDFs, publisher XML, HTML and supplementary files into checkable numbers — that is how the figures are obtained, not a check on them. The image-integrity screen is a pipeline of twelve discrete stages rather than a single check, so it has its own page. And the tortured-phrases screen is withheld while it remains quarantined behind a seed dictionary — it runs, but it is not yet calibrated enough to describe as a working check.

Each card below states the worst verdict its tool can deliver. That ceiling is a real limit, not a formality — most of these tools refuse rather than guess when a premise they depend on is missing, and several are capped below what their arithmetic would technically license because an innocent explanation always remains. Two are marked quarantined: they run, but the pipeline deliberately withholds their verdicts until their calibration criteria are met.

Granularity — the GRIM family10 tools

Integer data cannot produce just any summary statistic. Whole-number responses — Likert items, counts, binary outcomes — force means, percentages and standard deviations onto a discrete lattice fixed by the sample size and the reported precision. A value off that lattice was not computed from the data described. These are the toolkit's strongest checks: they need nothing but the numbers already printed, and a failure is arithmetic rather than opinion.

GRIM

grim_check

When it applies

A reported mean of data where every subject contributed a whole number, with the sample size known.

How it works

The mean of n integers must be some integer total divided by n. GRIM asks whether any integer total rounds to the mean exactly as printed. The tolerance scales with n rather than being a fixed half-unit, and both the round-half and truncation reporting conventions are tried before anything is flagged.

Inputs

mean
numberAs printed, not recomputed.
n
integer
decimals
integerPrinted decimals of the mean — “5.9” is 1, “5.90” is 2. There is no default: a wrong value accuses, so a missing one is refused.
scale_min / scale_maxoptional
integers
data_typeoptional
integer / continuous / unknownContinuous data is refused outright rather than judged.
discreteness_sourceoptional
stated / logically_necessary / assumedWho says each subject contributed a whole number. On an assumed premise a granularity failure caps at suspicious.

Output

Consistent, or a failure naming its mode — granularity or range violation, reported separately rather than fused. Only a failure resting on a stated or logically necessary premise reaches impossible; assumed bounds, assumed discreteness and truncation-only passes all cap at suspicious. Values read off a figure are refused, not judged.

Reference

Brown, N. J. L., & Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. Social Psychological and Personality Science 8(4), 363–369. 2017.

GRIM for percentages

grim_percentage

When it applies

A percentage that is a count's share of the sample — considerably more powerful than GRIM on a mean, because percentages carry more precision.

How it works

A share of a sample must be 100·k/n for some integer k, so at a given n only a finite set of percentages exists. Both round-half and truncation conventions are tried: SPSS and several journal styles truncate, and 9/93 = 9.6774 printed as “9.6%” is a truncating reporter rather than an impossibility.

Inputs

percentage
number
n
integerMust be the true denominator of THIS percentage, not the study total.
decimalsoptional
integer
data_typeoptional
stringDefaults to unknown deliberately, so an omission refuses rather than assumes.

Output

Whether the percentage is achievable at that sample size. Means, concentrations and ratios are refused — they are not constrained to multiples of 100/n. A truncation-only pass caps at suspicious.

Reference

Brown, N. J. L., & Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. Social Psychological and Personality Science 8(4), 363–369. 2017.

GRIM sweep

grim_sweep

When it applies

Several percentages that are supposed to share one denominator — a subgroup breakdown, or a row of a demographics table.

How it works

GRIM normally needs the sample size. Here the sample size is what is missing or disputed, so the logic runs backwards: for each reported value, enumerate every candidate n at which that value is achievable, then intersect those sets across all the values. What survives is the set of sample sizes that could have produced the whole row at once. An empty intersection is repeated under the truncation convention before any verdict is formed.

Inputs

values
list of numbers
n_max
integer
n_max_sourceoptional
stated / logically_necessary / assumedLoad-bearing, and it gates the verdict together with n_min: only a stated n_max plus an n_min reaching the floor licenses impossible.
n_minoptional
integer
decimalsoptional
integer

Output

The viable sample sizes per value and the joint intersection. A single surviving n is the recovered denominator; several mean the row does not pin it down. An empty intersection is NOT automatically a finding, and three outcomes are distinguished. If the values do share an n under TRUNCATION — which SPSS and several journal styles use — the verdict is consistent and it is recorded as a reporting convention, because a truncating reporter is not an impossibility. If no convention rescues them, impossible requires that the searched range provably bracket the true n: n_max must be the paper's own stated total and n_min must reach the floor. Otherwise the result is indeterminate, and says so — an empty intersection inside a range WE chose is a fact about the search, not about the paper.

Reference

Brown, N. J. L., & Heathers, J. A. J. The GRIM Test: A Simple Technique Detects Numerous Anomalies in the Reporting of Results in Psychology. Social Psychological and Personality Science 8(4), 363–369. 2017.

GRIMMER

grimmer_check

When it applies

A standard deviation reported beside a mean and sample size for integer data.

How it works

Standard deviations are granular for the same reason means are. GRIMMER enumerates the achievable pairs of sum and sum-of-squares behind the reported mean, subject to the parity constraint that the two must agree modulo 2, and asks whether any of them produces the reported SD at its printed precision.

Inputs

mean
number
sd
number
n
integer
decimals / decimals_sdoptional
integersPrinted precision of each. A wrong assumed precision accuses.
scale_min / scale_maxoptional
integers
discreteness_sourceoptional
stated / logically_necessary / assumed

Output

Consistent, or the constraint that failed. Only integer and parity failures license impossible; range-dependent failures on assumed bounds cap at suspicious, as do granularity modes on assumed discreteness. Both the sample and population SD conventions are tried, and a population-only match caps at suspicious.

DEBIT

debit_check

When it applies

A binary (0/1) variable reported as a mean and an SD — very common in Table 1 of a clinical paper.

How it works

For binary data the SD is not free: the mean fixes it exactly. DEBIT checks the reported SD against the value the mean and n force, testing both the sample (n−1) and population conventions before saying anything.

Inputs

mean
number (a proportion)
sd
number
n
integer
decimals_mean
integerPrinted decimals. A defaulted 2 dp falsely accused 73.6% of 1 dp rows in testing.
decimals_sd
integer
data_typeoptional
binary / proportionAnything else is refused rather than judged — a tool pointed at the wrong data must refuse, not accuse.

Output

Consistent, or an inconsistency with the SD the data would have to have. A row consistent under the population-SD convention is suspicious, never impossible — that convention is legitimate, and treating it as fraud produced 66 false accusations in a 450-case validation grid.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

DEBIT (whole table)

debit_check_batch

When it applies

An entire demographics table of binary variables, screened in one pass.

How it works

Runs the same arithmetic over every row, then judges the table rather than the row. Escalation requires a pattern of impossible rows; a lone hit stays suspicious.

Inputs

rows
table of rows (mean, SD, n, printed decimals)
data_typeoptional
binary / proportionRequired in practice — omitting it refuses the whole batch.
decimals_mean / decimals_sdoptional
integersSupply per row; the batch default of 2 dp falsely accuses 1 dp proportions.

Output

Every non-consistent row plus an aggregate verdict, and always the denominator — three suspicious rows out of 140 checks and three out of four are very different papers. A batch in which no row could be evaluated is indeterminate, never consistent.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

2×2 reconstruction

reconstruct_2x2

When it applies

A contingency result reported only as two column percentages and a total N, with the underlying counts withheld.

How it works

Enumerates every column split of the total and keeps the 2×2 tables whose cell counts round to both printed percentages. If a reported χ² is supplied, the candidate set is filtered by it as well.

Inputs

pct_col1 / pct_col2
numbers
n_total
integer
decimalsoptional
integer
reported_chi2optional
number

Output

The set of reconstructable tables, ordered by group-balance plausibility, or a proof that none exists. Impossibility is decided over every column split — the balance ratio only orders the survivors. A non-matching χ² yields indeterminate, never an accusation.

GRIM-U

grimu_check

When it applies

A p-value attributed to a Mann–Whitney U or Wilcoxon rank-sum test, with both group sizes known. Nothing else — t, F and χ² p-values are refused.

How it works

The U statistic takes only (half-)integer values between 0 and n₁·n₂, so at given group sizes only a finite set of p-values is achievable. The modelled set spans the normal approximation with and without continuity correction, plus the exact permutation p.

Inputs

n1 / n2
integers
reported_p
text, e.g. “0.043” or “<0.001”Passed as text so a threshold keeps its “<” — stripping it has turned an indeterminate result into a headline accusation.
test_typeoptional
mann_whitney / wilcoxonAny other test is refused, not scored.
one_sidedoptional
true / false

Output

Impossible below the U = 0 floor under every convention; suspicious inside a granularity gap; otherwise consistent. Known limitation, stated in the result: the tie-corrected variance that R and scipy use on tied data is not in the modelled set, so an impossible verdict on heavily tied data is unproven until checked again by hand.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

GRIM-U coexistence

grimu_coexistence

When it applies

Two or more nearly identical rank-sum p-values reported at the same group sizes — say 0.171 and 0.172.

How it works

Because the achievable p-values are a finite, unevenly spaced set, two p-values differing by less than the local spacing cannot both be real. This maps each reported value to the achievable U's and asks whether they can coexist.

Inputs

n1 / n2
integers
p_values
list of numbers
one_sidedoptional
true / falseForcing the two-sided set on one-tailed reports manufactures accusations.

Output

Whether the reported values can all arise at these group sizes. Carries the same tie-correction limitation as GRIM-U.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

SPRITE

sprite_reconstruct

When it applies

A mean and SD on a bounded integer scale, when you want to see what the underlying data could have looked like — or prove it could not have existed.

How it works

Searches for integer samples on the stated scale whose mean and SD round to the reported values, matching the SD at its reported precision rather than on exact variance. Alongside the search runs an analytic integer-and-parity screen that can prove no sample exists.

Inputs

mean
number
sd
number
n
integer
scale_min / scale_maxoptional
integers
decimals_mean / decimals_sdoptional
integers
discreteness_sourceoptional
stated / logically_necessary / assumed

Output

One of four outcomes, and only one of them is evidence: solutions found, no solution proven (from the analytic constraint), GRIM failure, or search exhausted. An exhausted search is a failed search, never an anomaly — conflating the two is how a reconstruction tool becomes an accusation generator. Recovered distributions with implausible shapes are returned for a human to judge.

Reference

Heathers, J. A., Anaya, J., van der Zee, T., & Brown, N. J. L. Recovering data from summary statistics: Sample Parameter Reconstruction via Iterative TEchniques (SPRITE). PeerJ Preprints 6, e26968v1. 2018.

Test-statistic recomputation10 tools

A test statistic, its degrees of freedom and its p-value are redundant with one another and with the group summaries they came from. Any one can be recomputed from the others, and the recomputed value has to agree with what was printed. These tools do that arithmetic — always against the interval the printed rounding allows, never against a point value, because matching a rounded input to an exact expectation is how a recomputation tool starts accusing honest papers.

statcheck

statcheck

When it applies

Any paper reporting results in APA format — a free, broad first pass over the whole text.

How it works

Extracts APA-formatted test statistics from the prose and recomputes each p-value from its statistic and degrees of freedom, flagging inconsistencies and, separately, decision errors where the recomputed p falls on the other side of the significance threshold.

Inputs

text
the paper's textA PDF path is accepted as a fallback route.

Output

Every extracted result with its reported and recomputed p, marked consistent, inconsistent or a decision error. Each p is recomputed assuming a two-tailed unadjusted test, and only APA grammar is visible to it — everything else is invisible, so a clean run is not a clean paper. Reporting inconsistencies are near-universal in honest work and must never support a misconduct verdict alone.

Reference

Nuijten, M. B., Hartgerink, C. H. J., van Assen, M. A. L. M., Epskamp, S., & Wicherts, J. M. The prevalence of statistical reporting errors in psychology (1985–2013). Behavior Research Methods 48(4), 1205–1226. 2016.

Independent-samples t-test

recalc_independent_t

When it applies

Two groups' means, SDs and sizes reported alongside a t or a p.

How it works

Recomputes t and p from the six group statistics and compares them against the interval the printed precision allows. Both Welch's and Student's pooled readings are accepted, and both tail conventions, unless the paper states which it ran.

Inputs

mean1, sd1, n1
numbers and an integer
mean2, sd2, n2
numbers and an integer
reported_t / reported_poptional
numbers
use_welchoptional
true / falseOmit when the paper does not say; both standard tests are then accepted.
p_tailoptional
one / two / unknown
decimals_mean / decimals_sd / decimals_poptional
integers

Output

The recomputed t and p with the achievable interval, and a verdict on the reported values. Convicting a pooled t against a Welch recomputation is a false accusation, which is why neither reading is assumed.

Within-subjects t-test

recalc_within_subjects_t

When it applies

A paired (pre/post) design reporting both timepoints' means and SDs plus a t or p, where the correlation between them is not stated.

How it works

A paired t implies a specific pre/post correlation, which can be recovered algebraically from the summary statistics. A recovered correlation outside [−1, 1] describes data that cannot exist. Because a thresholded p only lower-bounds t, the recovered r is then a lower bound too.

Inputs

mean_pre, sd_pre, mean_post, sd_post
numbers
n
integer
reported_p
text, e.g. “0.03” or “<.05”Passed as text so a threshold keeps its “<”. Stripping it once turned an indeterminate interval into a headline finding against a real paper.
p_is_thresholdoptional
true / false
p_tailoptional
one / two

Output

The recovered correlation and the whole interval attainable across the inputs' rounding box. Flagged only when that entire interval lies outside [−1, 1]; an exact-p impossibility caps at highly suspicious, because a reporting convention rather than fabrication is the usual cause.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

One-way ANOVA

recalc_anova

When it applies

A one-way ANOVA reported with its group means, SDs and sizes.

How it works

Rebuilds the between- and within-group mean squares from the group summaries, recomputes F and p, and compares them against the interval the printed rounding allows — rounding propagates multiplicatively into F, so a point comparison would be wrong.

Inputs

means / sds / ns
lists of numbers
reported_f / reported_poptional
numbers
decimals_mean / decimals_sd / decimals_poptional
integers

Output

Recomputed F and p with their achievable intervals. Assumes a classical fixed-effects one-way ANOVA rather than Welch's, with homogeneous variances and independent observations — all of which are stated in the result.

Chi-squared

recalc_chi_squared

When it applies

A contingency table whose counts are printed alongside a χ² or a p.

How it works

Recomputes χ² from the observed counts under both the Pearson and Yates conventions, and also computes Fisher's exact p for 2×2 tables. All three are tried before anything is flagged.

Inputs

observed
table of counts (rows × columns)
reported_chi2 / reported_poptional
numbers
use_yatesoptional
true / false

Output

Each convention's statistic and p against the reported values. The likelihood-ratio (G-test) convention is deliberately not tried and the result says so — a paper reporting G can mismatch entirely innocently.

F → p

recalc_f_to_p

When it applies

Any reported F with both degrees of freedom and a p — usable on ANOVAs of any shape, not just one-way.

How it works

The printed F is rounded, so it asserts a range of p rather than a value. This computes that range from the F's printed precision and its degrees of freedom, and checks the reported p against it.

Inputs

f_value
number
df1 / df2
integers
reported_poptional
number
decimals_f / decimals_poptional
integers

Output

The implied p-range and a verdict. Directional: only a p below the achievable range is flagged, since an under-claimed p is conservative reporting rather than an error.

Any statistic → its p-value

recalc_stat_to_p

When it applies

A single printed t, χ², z, r or F with a p beside it — the general-purpose version of the checks above.

How it works

Recomputes the p from the named distribution and degrees of freedom. Because the statistic is rounded it implies a p-interval, and the reported p is judged against that interval rather than a point.

Inputs

statistic
number
kind
t / chi2 / z / r / F
df / df2 / noptional
integersWhichever the named test needs; a missing one is named in the refusal.
reported_poptional
number
tailsoptional
1 or 2Ignored for χ² and F.

Output

The implied p-interval and a directional verdict — only an over-claim is flagged. Capped at suspicious, tighter than its class allows: a one-tailed test, a different df convention or an unshown correction would each explain a disagreement. A missing df or n is reported as an incompletely reported test, which is a property of the paper.

Regression coefficient

check_regression

When it applies

A regression table printing a coefficient, its standard error, and a t or p.

How it works

Checks that the reported t equals B divided by SE, judged against the interval reachable from the printed precision of B and SE, and that the p is that t's tail probability at the given degrees of freedom. Both sign conventions (signed t and |t|) and both tail conventions are accepted first.

Inputs

b
number
se
number
reported_t / reported_poptional
numbers
dfoptional
integerWithout it only the p check is skipped; t = B/SE still runs.
p_is_thresholdoptional
true / falseTrue when the paper printed a bound such as “P < .001” — a float cannot carry that difference.

Output

The implied t, its achievable interval, and verdicts on the reported t and p separately.

Regression table (whole)

check_regression_batch

When it applies

Every coefficient in a regression table, in one call.

How it works

Runs the coefficient check per row. Where the degrees of freedom were not reported it re-runs at df of 10, 30, 100 and 100,000; a verdict that changes across that range is returned as indeterminate rather than resolved by a guess.

Inputs

rows
table of rows (b, se, reported t/p, df, printed decimals, label)

Output

Per-row results plus an aggregate set by the worst row. A batch in which no row could be evaluated is indeterminate, never consistent — and the denominator is always reported alongside the hits.

STALT — hidden p-values

stalt_check

When it applies

A paper reporting only “p < 0.05” where the group statistics permit an exact recalculation.

How it works

Compares the recomputed exact p against the threshold the paper printed. The threshold is read as an upper-bound claim, so any p at or below it is consistent; what the check surfaces is a p many orders of magnitude below the stated bound — information the paper had and did not report.

Inputs

calculated_p
numberMust come from the same test the printed threshold describes.
reported_thresholdoptional
text, e.g. “<0.05”

Output

The size of the overshoot and a severity that scales with it. A factor-of-two overshoot is tolerated as a one- versus two-tailed artefact rather than flagged.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

Table arithmetic & internal consistency10 tools

Papers state the same quantity more than once, in different forms, and the forms have to agree. A count and its percentage share a denominator; subgroups sum to their total; an estimate sits at the centre of its own confidence interval; a standard deviation cannot exceed what the variable's range permits. None of this needs raw data — only that the numbers printed in one table be consistent with the numbers printed beside them.

Whole-table check

check_table

When it applies

Any table — this is the entry point, and it replaces calling the per-cell tools one at a time.

How it works

A router. Given typed cells, the variables you declared and a binding from row labels to variables, it determines which checks apply from cell kind crossed with variable family and runs all of them. Which checks apply is derived, not chosen. Declaring a variable also resolves cells that syntax alone must refuse: “48.90 (14.46)” under a bare label is ambiguous until the variable is declared continuous.

Inputs

cells
typed table cells
variablesoptional
list of variable declarationsDeclare from the paper, never from a guess — a fabricated bound inverts a check rather than weakening it.
bindingsoptional
map of row label → variable
provenanceoptional
printed_table / measured_raster / ocr_table / …Figure-read and OCR'd values make the exact-arithmetic checks refuse automatically.

Output

Every non-consistent row in full, a count of the consistent ones, and a machine-generated list giving the reason each check was refused — so coverage does not depend on anyone remembering. It also returns a declaration template naming exactly which variables are blocking checks it could otherwise run. The impossible ceiling is the router's, not every row's: read each row's own tool to know what its verdict rests on.

Count against percentage

check_count_percent

When it applies

Any “n (%)” cell — which is most of Table 1 in a clinical paper. Use it whenever the count is printed; it is strictly stronger than GRIM on a percentage.

How it works

Divides the printed count by the printed denominator and checks the result against the printed percentage, at the rounding interval the percentage's own precision asserts.

Inputs

count
integer
percentage
number
n
integer
decimalsoptional
integerWithout it the check returns indeterminate rather than assuming a precision.

Output

A verdict plus the implied denominator — and across a table the implied denominators are the real evidence, because a quantity that can only have one value should not imply several. Capped at suspicious, tighter than its class allows: row, column and subgroup bases legitimately coexist in one table, and no arithmetic disagreement rules that out. A count exceeding its denominator is reported as a wrong pairing on our side, not a defect in the paper.

Ratio check

check_ratio

When it applies

A printed ratio shown with its own numerator and denominator — cost-effectiveness ratios, rates, “X per Y”.

How it works

Computes the quotient at the corners of the rounding box implied by all three printed precisions, and asks whether the printed ratio falls inside it.

Inputs

numerator
number
denominator
number
reported_ratio
number
decimals_numerator / decimals_denominator / decimals_ratiooptional
integersAll three are needed; without them the check refuses, because a guessed precision either clears every ratio or accuses on rounding alone.

Output

A verdict plus the implied numerator and denominator. Capped at suspicious despite testing exact arithmetic: papers routinely divide unrounded intermediates they never show, or a discounted or subgroup quantity. Across a table, several rows implying different values for a quantity that can only have one is the finding — not any single row. A denominator whose rounding interval spans zero makes the quotient unbounded, and the check refuses rather than compares.

Summation check

check_summation

When it applies

Values that should add to a stated total — subgroup counts to N, percentages to 100, cost components to a total cost.

How it works

Sums the addends and compares against the stated total within a window derived from the printed decimals of every addend and of the total, so rounding is accounted for rather than assumed away.

Inputs

values
list of numbers
expected_sum
number
decimalsoptional
list of integersPrinted decimals per addend. With neither this nor an explicit tolerance the check refuses rather than guessing.
toleranceoptional
numberCount columns that must hold exactly take 0; percentage columns legitimately misbalance by rounding.

Output

The computed sum, the permitted window and the discrepancy. The values are assumed to be exhaustive, non-overlapping addends — an omitted row or an overlapping category mimics a mismatch, and the result says so.

Summation (whole table)

check_summation_batch

When it applies

Every summation in a table, in one call.

How it works

Runs the same check per row, with per-row printed decimals, and sets the aggregate verdict from the worst row.

Inputs

rows
table of rows (values, expected total, printed decimals)
toleranceoptional
numberA batch-wide override.

Output

Per-row results and an aggregate. A row with no printed precision is refused rather than guessed, and a batch in which nothing could be evaluated is indeterminate, never consistent.

Value against its logical range

check_statistic_bounds

When it applies

A value that appears to fall outside what its measure permits by definition — a correlation above 1, a p-value above 1, a percentage above 100.

How it works

Only bounds that follow from a measure's definition are known to it — a correlation is bounded by Cauchy–Schwarz, a p-value by the definition of a probability. A merely expected range is an assumption, so unknown measures are refused rather than judged against our expectations. Crucially, it treats a violation as evidence about our extraction before evidence about the paper.

Inputs

value
number
measure
correlation / p_value / percentage / …Anything without a definitional bound is refused, not guessed at.
extraction_verifiedoptional
true / falseSet only after re-reading the value in the paper and confirming both the digits and the column. This is what lifts the refusal to impossible.

Output

By default, indeterminate with the failure mode “extraction suspect”, naming the decimal-point slip or wrong-column read that would explain it — because gross violations are uncommon even in fraudulent papers, since fabricators produce plausible numbers. Verification is the price of the accusation. Even when verified, the tool names the innocent readings, because a confirmed impossible value still does not establish who produced it.

SD range check

check_sd_range

When it applies

A standard deviation for a variable whose attainable range is known.

How it works

The largest SD a bounded sample can have is fixed by its range — (range/2)·√(n/(n−1)) for the sample SD. A reported SD above that ceiling describes data that cannot exist. With n unknown the ceiling is taken at its weakest over all n; supplying n tightens it.

Inputs

sd
number
min_val / max_val
numbersMust be the variable's true attainable bounds.
noptional
integer
bound_sourceoptional
stated / logically_necessary / assumedLoad-bearing: only a range the paper states, or one that is logically necessary, licenses impossible. An assumed range caps at highly suspicious.

Output

The ceiling, the reported SD and the verdict. A negative SD is bound-free and impossible whatever else is unknown.

SD-or-SE adjudication

check_sd_or_se

When it applies

A dispersion value where it is unclear whether the paper reported a standard deviation or a standard error.

How it works

SD and SE are linked by SE = SD/√n, so both readings can be tested against a plausibility ceiling implied by the variable's range and the reading that fits adjudicated. Under the SE reading the rounding interval stretches by √n, which is accounted for.

Inputs

reported_value
number
n
integerMust be the cell's true denominator.
reported_asoptional
sd / se / unknown
variable_range_min / variable_range_maxoptional
numbers
bound_sourceoptional
stated / logically_necessary / assumed

Output

Which reading the value must be, or that neither fits. A “neither fits” verdict rests entirely on the plausibility ceiling, so on an assumed bound it caps at suspicious. The result carries the base rate: mislabelling SE as SD is the norm rather than an outlier — Olsen found 35 of 88 studies in one journal doing it — and it is usually a statistical error, not a sign of misconduct.

Reference

Olsen, C. H. Review of the Use of Statistics in Infection andImmunity. Infection and Immunity 71(12), 6689–6692. 2003.

Estimate against its own interval

check_estimate_ci

When it applies

A point estimate printed with a confidence interval, and often a p-value too — odds ratios, hazard ratios, mean differences.

How it works

Checks ordering, containment and the midpoint on the correct scale — linear for a difference, logarithmic for a ratio measure — and recovers the p implied by the interval to compare against the reported one. Alternative confidence levels are tried before any “wrong p” is reported.

Inputs

estimate
number
ci_low / ci_high
numbers
measureoptional
OR / RR / HR / IRR / differenceDecides whether symmetry is checked on the linear or the log scale.
reported_poptional
number
reported_p_operatoroptional
=, <, ≤, >, ≥Pass “<” for “P < .001” — a threshold is an interval claim, not a value.

Output

Each sub-check separately with its verdict. Exact, profile and bootstrap intervals are legitimately asymmetric and can fail the midpoint check innocently, which the result states. A p that cannot be placed against alpha at its own printed precision is left unclassified rather than judged.

Continuous → categorical plausibility

check_proportion_from_normal

When it applies

A paper reporting both a continuous variable's mean and SD and the proportion of participants above or below some threshold on it.

How it works

Given the mean, SD and n, the share of the sample past a threshold is constrained. The headline test assumes the raw variable is normal and uses a binomial test on the count at the threshold; a second, distribution-free test uses Cantelli's inequality, which holds for every distribution.

Inputs

reported_proportion
number
mean / sd
numbers
n
integer
threshold
number
directionoptional
above / below

Output

Both tests separately. The normal-assumption result caps at suspicious, because skew explains a mismatch innocently; only a violation of the distribution-free Cantelli bound reaches highly suspicious. The proportion and the mean/SD/n must describe the same sample.

Similarity & duplication5 tools

Real data is noisy in ways people reliably fail to imitate, and it is never noisy twice in the same way. Randomised groups differ at baseline by chance; sample standard deviations scatter by a knowable amount; two independent samples do not produce the same mean and SD. Fabricators add noise to the means, because everyone knows means vary, and forget that the dispersion measures need their own — and when a table or a spreadsheet is filled by copying, the duplicate values survive in the published numbers. The first three tools here test for data that is too tidy; the last two are copy-paste detectors, working on printed summary statistics and on participant-level raw data respectively. None of them can ever return “impossible”: no arrangement of numbers is forbidden by arithmetic, only improbable.

Carlisle–Stouffer–Fisher test

csf_test

When it applies

A randomised trial's Table 1, where baseline characteristics are compared between arms.

How it works

Under simple randomisation, baseline p-values are uniform and independent. Combining them by Stouffer's or Fisher's method gives a one-sided test for EXCESS similarity. A very small result means the arms are more alike than randomisation can plausibly produce — the signature of a trial whose allocation never happened.

Inputs

p_values
list of numbersAt least three usable values are needed.
designoptional
simple / cluster / stratified / crossover / …Cluster, multi-site, stepped-wedge, crossover, matched, stratified and minimised designs force baseline balance by construction and are refused outright. An unstated design caps the verdict.
correlationoptional
numberBaseline rows are correlated; with none supplied an equicorrelated null is read at a conservative ρ = 0.5.

Output

A one-sided similarity p-value with a sensitivity profile across correlation assumptions. High baseline p-values are the red flag here, never reassurance — a Table 1 where every p is 0.93, 0.99, 0.99 is the thing this test exists to catch.

Bayesian Table-1 dispersion

bayesian_table1_dispersion

When it applies

The same Table 1 as the CSF test — complementary to it, not preferred over it, and handling mixed continuous and categorical rows.

How it works

Turns each baseline row into a p-value and then a z-score, and estimates a precision multiplier τ̂ = √(1/mean(z²)). Above 1 indicates under-dispersion — differences systematically smaller than chance allows; below 1, over-dispersion and possible randomisation failure. Rows are treated as equicorrelated, discounting a table of any size to about two effective degrees of freedom, so a large table is not thereby more accusable.

Inputs

rows
list of baseline rows — means/SDs/ns, or categorical counts
designoptional
stringSame refusals as the CSF test; an unstated design caps the verdict below highly suspicious.
correlationoptional
numberPass only if genuinely known or estimable.

Output

τ̂ with a confidence interval that inverts the exact chi-square null, plus a sensitivity profile. An elevated τ̂ whose interval does not exclude 1 is reported indeterminate, never consistent — too few effective degrees of freedom to convict OR to clear. Escalation additionally requires the CSF test on the same p-values to corroborate; CSF finding nothing vetoes it.

Reference

Bolland, M. J., Gamble, G. D., Avenell, A., Grey, A., & Lumley, T. Baseline P value distributions in randomized trials were uniform for continuous but not categorical variables. Journal of Clinical Epidemiology 112, 67–76. 2019.

Variance dispersion

check_variance_dispersion

When it applies

Three or more groups whose standard deviations are reported — are the SDs too similar to one another?

How it works

Sample variances are themselves random: run the same experiment on several groups and the SDs scatter by a knowable amount. Standardising by the within-group mean square leaves only relative spread, whose null is bootstrapped from the chi-square distribution. Dispersion is measured as the range of the z-scores rather than their SD, which the source study found more robust to heterogeneity.

Inputs

sds / ns
lists of numbersAt least three groups.
sd_texts
list of the SDs as printedPrecision is read from the literal; rounding coarser than a quarter of the SD is refused, because coarse rounding drives honest SDs onto identical values.
homogeneous_variancesoptional
true / falseRequired — a violated assumption blinds the test rather than making it accusatory, so it gates the consistent verdict too.
designoptional
stringDesigns that constrain variances by construction are refused.

Output

The observed dispersion against its bootstrapped null, capped at suspicious — over-similar SDs are improbable, never impossible. Every result sets a reference-required flag: the validating study's near-perfect discrimination came from comparing genuine against fabricated sets, while both looked significant against the theoretical null in absolute terms. This is a relative verdict, and the result says so.

References

Hartgerink, C. H. J., Voelkel, J. G., Wicherts, J. M., & van Assen, M. A. L. M. Detection of data fabrication using statistical tools. PsyArXiv . 2019.

Simonsohn, U. Just Post It: The Lesson From Two Cases of Fabricated Data Detected by Statistics Alone. Psychological Science 24(10), 1875–1888. 2013.

Summary-statistic copy-paste

detect_duplication

When it applies

Across a paper's tables and studies, wherever summary statistics are supposed to describe independent samples — and within a single table, where a filled-down row or column leaves the same values behind.

How it works

Hashes every (mean, SD) cell at its PRINTED precision and looks for three signatures. Across blocks the paper calls independent samples, it counts identical pairs and prices them against a Poisson chance model whose value spans are inferred from the data itself rather than assumed. Within a single table, it looks for a whole row or column of cells repeated under a different label — the fill-down signature. And it reports pairs identical after a single digit transposition, as a weak pointer only.

Inputs

blocks
labelled blocks of cells, each with a sample identityBlocks are assumed independent unless they share a sample id — Tables 1, 2 and 3 of a cohort paper are usually subsets of each other.
decimals_mean / decimals_sdoptional
integersThe printed precision cells are hashed at. Cells may override per-cell.
near_matchoptional
true / falseAlso report pairs identical after a digit transposition. Always indeterminate — a pointer, never a finding.

Output

Each duplicate group with its kind, the matched values, and a multiplicity-corrected chance probability. Severity keys on the number of identical pairs — three escalates — and several guards pull it back down. Blocks sharing a sample id, or declared non-independent, drop two tiers. Round values (integers, halves, multiples of 0.05) collide far more often and drop one. A lone match is re-checked against its own chance probability and demoted if a coincidence was likely. Within a table, integer count cells with no SD never count toward severity at all: blocked allocation and zero attrition make arm sizes identical BY DESIGN. Capped at highly suspicious — duplication is improbable, never impossible.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

Raw-data copy-paste

detect_raw_duplication

When it applies

A participant-level supplement — a spreadsheet of rows, one per subject. The strongest evidence available, when the data exists.

How it works

Takes a file path and loads the grid itself — a raw sheet is thousands of values and no model can retype it — then looks for three signatures: duplicated participant rows within and across sheets, vertical runs of values reappearing in order, and improbably specific recurring numbers. Matching is exact after float canonicalisation; there is no tolerance anywhere, because a tolerance manufactures matches between genuinely different measurements. Evidence is weighted by the INFORMATION CONTENT of each number rather than by match count: 123.46 recurring means something, 100 recurring means nothing, and a value shared by twenty rows is a site-level covariate rather than twenty copy-pastes.

Inputs

path
spreadsheet file (.xlsx, .csv, .tsv, .docx)
sheetoptional
worksheet name
exclude_columnsoptional
list of column namesColumns shared by design — IDs, site codes, coordinates, plot-level covariates — must be excluded by the caller or they duplicate legitimately.
data_contextoptional
raw_participant_level / pooled_or_meta_analysis / unknownA pooled or meta-analytic file duplicates rows by construction; findings are then capped at indeterminate. That is a refusal, not a finding.

Output

Each signature with its own chance model — a Poisson null over same-precision bins for repeated values, an empirical collision null for matching runs — so a large sheet is not flagged merely for being large. Row-pair evidence has no chance model and the result says so plainly; it rests on information content alone and is the weaker limb. Capped at highly suspicious.

Reference

Heathers, James An Introduction to Forensic Metascience. forensicmetascience.com . 2025.

Design, power & inference5 tools

Not every problem in a paper is a wrong number. A stated power analysis can be recomputed against the sample actually recruited; a family of tests has an arithmetic false-positive rate whatever the paper says about it; a “no difference” conclusion can be checked against the effect its own confidence interval still permits. These tools audit the inferences rather than the arithmetic — and several of them are calculators that render no verdict at all.

Power-analysis verification

verify_power_analysis

When it applies

A paper stating an a-priori power analysis — the effect size, alpha and target power it planned around.

How it works

Recomputes the required sample size using exact noncentral distributions and compares it against the N actually recruited, with a small tolerance absorbing documented differences between G*Power, pwr and PASS.

Inputs

test_type
string
effect_size / alpha / target_poweroptional
numbers
n_reportedoptional
integer
tailsoptional
1 or 2As stated by the paper. Left unset, both readings are tried and no accusation is made when the one-tailed reading reconciles.
allocation_ratiooptional
numberAnything other than 1.0 is refused — the recomputation is 1:1, and grading a k:1 design against it can accuse a correct calculation.

Output

The required N against the recruited N. Indeterminate unless every input was stated — which is the modal honest outcome and is itself worth reporting: roughly 80% of published power analyses cannot be checked at all, because the inputs are not there. Over-recruiting is treated as correct practice, never a discrepancy. It cannot detect an analysis that reproduces but powers the wrong hypothesis.

Reference

Bakker, M., Veldkamp, C. L. S., van Assen, M. A. L. M., Crompvoets, E. A. V., Ong, H. H., Nosek, B. A., Soderberg, C. K., Mellor, D., & Wicherts, J. M. Ensuring the quality and specificity of preregistrations. PLOS Biology 18(12), e3000937. 2020.

Required sample size

required_n

When it applies

To ask what sample a design would have needed — as context for a study, not as a charge against it.

How it works

A pure calculator: the smallest n reaching a target power for a named test at a given effect size and alpha, via exact noncentral distributions. Supplying the sample actually used adds a design-sensitivity comparison — the effect the study could have detected.

Inputs

test_type
string
effect_size
number
alpha / poweroptional
numbers
actual_noptional
integerAdds the design-sensitivity block; without it the answer is a bare sample-size calculation.

Output

The required n, and optionally the minimum detectable effect. It renders no verdict about any paper: a small study is a fact about sensitivity, not a defect.

Multiple-comparisons arithmetic

multiplicity_report

When it applies

A paper running many tests and treating each at α = 0.05 — or claiming a correction that can be checked.

How it works

For a declared family of tests, computes the family-wise error rate and the thresholds four standard corrections would impose: Bonferroni, Holm, Benjamini–Hochberg and Benjamini–Yekutieli. It then counts how many reported claims survive each.

Inputs

tests
list of tests (p-value, whether a correction was applied)
family_label
textDescribe the family you chose. Family definition is a judgement the tool cannot make, so there is no default — an omitted label is refused, not filled in.
alphaoptional
number

Output

The family-wise error rate, the four thresholds, and floor and ceiling counts of surviving claims. Everything is conditional on the declared family — the same p-values give different answers under different families — and the survivor count is a floor, because undisclosed analytic flexibility means the true comparison count can exceed the reported one. Deterministic arithmetic; it accuses no one.

References

Benjamini, Y., & Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B 57(1), 289–300. 1995.

Benjamini, Y., & Yekutieli, D. The control of the false discovery rate in multiple testing under dependency. The Annals of Statistics 29(4), 1165–1188. 2001.

Absence of evidence

absence_of_evidence

When it applies

A paper concluding “no effect” or “no difference” from a non-significant result.

How it works

“Not significant” is not “no effect”. This reports the largest effect the paper's own confidence interval still permits, and judges the CLAIM made about the interval rather than the data. If no interval is printed, one is derived from the coefficient and standard error and marked as derived so it never passes as the paper's.

Inputs

ci_low / ci_highoptional
numbersOr pass b and se instead and an interval is derived.
measureoptional
OR / RR / HR / IRR / differenceRequired in practice: the null is 1 for a ratio measure and 0 for a difference, and without it the tool declines. Omitting it once inverted a verdict — the interval [0.85, 1.42] was reported as excluding the null.
sesoioptional
numberSmallest effect size of interest. Without one, equivalence is described but not judged — that cannot be claimed against an unstated threshold.
dfoptional
integerWithout it the normal distribution is used and the result says so.

Output

The interval, the largest effect it still admits, and whether the stated conclusion survives it. Declines rather than guesses when the null cannot be located or no threshold was supplied.

Effect-size plausibility

check_effect_plausibility

When it applies

Any two-group comparison — and notably the one check that still works on figure-derived data, so a paper reporting its outcome only as a bar chart becomes checkable here.

How it works

Computes Cohen's d and Hedges' g with a confidence interval from the six group statistics. Passing the paper's own reported d alongside them checks internal consistency instead: does the printed effect size match the one its own means, SDs and ns imply? That needs no comparison class and is what catches a mistyped effect size.

Inputs

mean1, sd1, n1, mean2, sd2, n2optional
numbers and integersOr pass d directly.
reported_doptional
numberTriggers the internal-consistency check.
fieldoptional
stringUsually omitted. Adds field benchmarks — but only social psychology's percentiles trace to a published table; the rest are uncited estimates and say so.
provenanceoptional
printed_text / figure_pixels

Output

The effect size with its interval — reporting the value is the deliverable. Capped at suspicious: an effect can be improbable but never arithmetically impossible, and an unusually large one is informative, not evidence of misconduct. Before treating a large effect as a finding, three ordinary causes are named in the result: SEM read as SD (which inflates d by √n), a waitlist rather than active comparator, and a units mismatch. Field benchmarks are distributions of published effects, themselves inflated by publication bias.

Reference

Lovakov, A., & Agadullina, E. R. Empirically derived guidelines for effect size interpretation in social psychology. European Journal of Social Psychology 51(3), 485–504. 2021.

Hand-calculation signatures & heuristics4 tools

Some anomalies are not impossible numbers but improbable patterns: a test statistic that exactly equals the value you would get by hand from the rounded table, sample sizes too round to be real recruitment, or a comparison reported so coarsely that no one can tell what was computed. These are triage signals. They escalate alongside stronger findings and are capped so they cannot convict on their own.

RIVETS (t-test)

rivets_independent_t

When it applies

A t-test reported beside the group summaries it supposedly came from — the precision-aware replacement for a plain recomputation.

How it works

Real data has more precision than the table prints, so a t computed from actual data almost never lands exactly on the value you get from the rounded summaries. RIVETS maps the t and p reachable across the inputs' rounding intervals by Monte Carlo, and measures how often an exact match occurs. A point match plus a rare hit rate is the signature of a statistic computed from the table rather than from data.

Inputs

mean1, sd1, n1, mean2, sd2, n2
numbers and integers
reported_t / reported_poptional
numbers
decimals_mean / decimals_sd / decimals_t / decimals_poptional
integersMust be supplied; the defaults assume more precision than most papers report.
p_is_thresholdoptional
true / falsePass a bound p as text, never as a float, or it reads as an unreachable exact value.

Output

The achievable t and p intervals, the exact-match hit rate, and a verdict. Both Welch and Student readings and both tail conventions are accepted unless the paper states them.

Reference

Brown, N. J. L., & Heathers, J. Rounded Input Variables, Exact Test Statistics (RIVETS). PsyArXiv . 2019.

RIVETS (ANOVA)

rivets_anova

When it applies

The same detector applied to a reported F. Preferred over the plain ANOVA recomputation, which is precision-blind.

How it works

Monte Carlo over the truncation interval of every printed input, producing the reachable F and p ranges plus the exact-match hit rates.

Inputs

means / sds / ns
lists of numbers
reported_f / reported_poptional
numbers
decimals_mean / decimals_sd / decimals_f / decimals_poptional
integers
p_is_thresholdoptional
true / false

Output

Reachable ranges, hit rates and a verdict. A reported p outside the achievable range is flagged at the same severity as an out-of-range F. This tool owns the accusation-grade reported-p check for ANOVAs.

Reference

Brown, N. J. L., & Heathers, J. Rounded Input Variables, Exact Test Statistics (RIVETS). PsyArXiv . 2019.

Round-N flag

flag_round_n

When it applies

A study reporting several distinct sample sizes that are all suspiciously round.

How it works

A base-rate heuristic. Fabricated papers over-represent round Ns — but so do honest recruitment targets, which is why it needs at least three reported sizes with at least two distinct before it says anything at all.

Inputs

ns
list of integers
moderate_modulus / strong_modulusoptional
integers

Output

A triage flag capped at suspicious — a weak signal that escalates only alongside stronger findings, and never a finding on its own.

Uninformative-statistic flag

flag_uninformative_stat

When it applies

A comparison reported so coarsely that the printed precision cannot pin down the test outcome.

How it works

Models the comparison as a one-way ANOVA over the group summaries and computes the whole p-range the printed precision permits — by Monte Carlo plus an exact corner search, because the extreme of F lives on a vertex that interior sampling reaches with probability zero. It then asks whether that range straddles the significance threshold.

Inputs

means / sds / ns
lists of numbers
reported_poptional
number
decimals_mean / decimals_sd / decimals_poptional
integers
significance_thresholdoptional
number

Output

The analytic p-range, with flags for a range straddling alpha, an implausibly wide range, or a reported p outside it. Capped at suspicious with the one-way-ANOVA assumption disclosed; the RIVETS ANOVA check owns anything accusation-grade.

Cross-checks against R2 tools

Two tools exist only to check our own arithmetic against the canonical R implementations. They render no verdict about any paper: a disagreement between engines means OUR tooling is wrong, and the finding is withheld rather than reported. An engine that cannot run is never a disagreement — the wrappers distinguish “package missing”, “R absent” and “R errored”, because a missing package once silently capped every GRIM finding in a whole corpus.

GRIM ground truth (R scrutiny)

r_grim_ground_truth

When it applies

Before publishing any GRIM result at impossible — the prompts require this cross-check first.

How it works

Shells out to R and runs the same test through the scrutiny package, then compares verdicts.

Inputs

mean
number
n
integer
itemsoptional
integer
decimalsoptional
integer

Output

R's verdict beside ours. A disagreement is a bug on our side and is never evidence about the paper. If R or the package is unavailable the result says which — it does not report a disagreement.

SPRITE ground truth (R rsprite2)

r_sprite_ground_truth

When it applies

To confirm a SPRITE no-solution result before it is relied on.

How it works

Runs the reconstruction through R's rsprite2 and reports whether it proves infeasibility. Note that rsprite2 infers decimal precision from the literal it is passed, so the precision must be pinned explicitly or the two engines are silently answering different questions.

Inputs

mean / sd
numbers
n
integer
scale_min / scale_maxoptional
integers
decimalsoptional
integerPin this explicitly — rsprite2 reads “4.0” as zero decimals.
seedoptional
integerVary it to test whether a no-solution result is robust.

Output

rsprite2's samples, or its proof that none exists. A disagreement with our implementation is a bug in our tooling and the finding must not be reported.

Reference

Heathers, J. A., Anaya, J., van der Zee, T., & Brown, N. J. L. Recovering data from summary statistics: Sample Parameter Reconstruction via Iterative TEchniques (SPRITE). PeerJ Preprints 6, e26968v1. 2018.

The toolkit as a whole follows James Heathers’ An Introduction to Forensic Metascience. Individual methods are credited on their own cards. See also the findings the toolkit has reported.