Checking for bias by comparing with "random sampled" replication initiatives
Computed from database version 2026_08_01 (8,445 replication pairs). Rates on this page use the site-wide success-rate definition — success / (success + failure + reversal), with inconclusive and unrecorded outcomes excluded from the denominator — the same rule every replications database page applies. Every figure below is regenerated by scripts/build_bias_check.py; re-run it whenever a new database version lands.
Most of the replications database is harvested from the published literature: we find papers that describe themselves as replications and extract what they report. That method has an obvious weakness. Nobody publishes a replication for a random reason, and a literature we assemble by searching it is not a sample of anything in particular. If replication attempts reach print mainly when they are interesting — a famous effect knocked down, or a contested one vindicated — then a success rate computed over that corpus tells us about publishing incentives as much as about science.
A handful of replication initiatives give us a way to test that. Projects like the Reproducibility Project, DARPA SCORE and the Brazilian Reproducibility Initiative each defined a sampling frame in advance — specific journals, specific years — and replicated what fell inside it, regardless of whether the result would be interesting. If our harvested corpus is biased, it should look systematically different from these pre-specified samples. This page runs that comparison.
The crude comparison is misleading
It is tempting to simply split the corpus into "coordinated replication initiatives" on one side and "ad-hoc replications harvested from the literature" on the other. Done that way — every row belonging to a coordinated initiative on one side (1,439 rows, 40%), every loose literature-harvested row on the other (6,996 rows, 59%) — it looks like the ad-hoc literature is 19 points rosier, as if published replications skew positive.
That comparison is misleading, and the real signal is more interesting. The "initiatives" bucket mixes two opposite selection strategies, and pulling them apart accounts for most of the gap:
| Selection strategy | # findings | Success rate |
|---|---|---|
| Representative / systematically-sampled initiatives (RP:P, DARPA SCORE, LOOPR, X-Phi, Cancer Biology, Brazilian Reproducibility Initiative, …) | 916 | 52% |
| Targeted mega-replications of famous / contested findings (Many Labs, Registered Replication Reports) | 350 | 11% |
| Ad-hoc / literature-harvested replications | 5,731 | 59% |
Counts are replication findings (many initiatives replicate several findings per paper) with a definitive outcome.
Conclusion: most of the "coordinated initiatives look worse" effect is not about coordination at all — it is the targeted mega-replications, which deliberately go after a handful of celebrated, hotly-contested effects and find that only ~11% hold up (Many Labs 3 6%, Many Labs 4 6%, Many Labs 5 10%, Registered Replication Reports 11%). Broad-sampling projects score far higher (Life Outcomes of Personality Replication Project 86%, Experimental Philosophy Reproducibility Project 78%, L2 Replication 71%, DARPA SCORE 53%). Separating the two buckets shrinks the 19-point gap to 7 points — representative-sampled 52% vs. literature-harvested 59%.
That remaining 7 points looks real rather than noise — a two-proportion test gives p < 0.0001 — though that test overstates the precision, because the initiative rows cluster within papers and projects while the harvested rows do not (see the third caution below). Treat the direction as solid and the interval as softer than the p-value implies: these two populations are not interchangeable. The gap is nonetheless largely a design difference rather than a quality difference: the representative initiatives are 75% direct replications, while the harvested rows are only 13% direct and mostly looser "close" variants — and design strongly predicts outcome (direct 42%, close extensions 66%). Compare the two on direct replications alone, as Table A below does, and the harvested advantage reverses in psychology — the one field with a large sample on both sides — though it persists elsewhere.
The load-bearing variable is what gets picked for replication, and how faithfully it is re-run — not whether the replication was part of an organized initiative. Hand-picking a famous, surprising claim — exactly the kind of selection that also drives which findings get harvested from the literature — remains one of the strongest predictors of failure in the whole database. (Caveat: the literature-harvested 59% is the least-validated number in this corpus — roughly two-thirds AI-curated with heterogeneous "success" definitions — so read it as suggestive, not settled.)
Within disciplines, the comparison depends on design
Breaking representative-sampled and literature-harvested replications down by field (restricted to cells with enough rows to be worth reading) shows this is not an artifact of psychology dominating the corpus — but it also shows the fields do not all point the same way.
The comparison is shown twice, because design leniency differs sharply between the two columns: the initiatives are 75% direct replications, while the harvested rows are only 13% direct and mostly looser "close" variants — and design predicts outcome (direct 42%, close extension 66%). The first table therefore compares like with like on design; the second uses everything we have harvested.
Table A — initiatives vs. our direct replications only
| Discipline | "Representative-sampled" initiatives | Literature-harvested (direct only) |
|---|---|---|
| Psychology | 58% (n=551) | 49% (n=302) |
| Biomedical | 31% (n=227) | 38% (n=82) |
| Economics | 59% (n=29) | 53% (n=17) |
| Education | 75% (n=16) | 90% (n=21) |
Table B — initiatives vs. all harvested replication types
| Discipline | "Representative-sampled" initiatives | Literature-harvested (all types) |
|---|---|---|
| Psychology | 58% (n=551) | 59% (n=2,770) |
| Biomedical | 31% (n=227) | 50% (n=361) |
| Linguistics | 66% (n=59) | 64% (n=121) |
| Economics | 59% (n=29) | 57% (n=315) |
| Education | 75% (n=16) | 84% (n=131) |
| Political science | 25% (n=4) | 58% (n=72) |
| Business & management | 50% (n=4) | 67% (n=49) |
| Sociology | 50% (n=2) | 50% (n=50) |
Cells show the site-wide success rate — success / (success + failure + reversal), inconclusive and unrecorded outcomes excluded — and n is that denominator. Cells with n < 20 carry no information — read them as blank, not as evidence. The initiative column is identical in both tables; only the harvested column changes, so a discipline can appear in one table and not the other. Table B lists every field where both columns are populated and the harvested side reaches n ≥ 20, which is why its lower half is thinner than its upper half: education, political science, business & management and sociology have initiative columns of n = 16, 4, 4 and 2 and are included for completeness, not as evidence — read those four cells as blank. Three of them rest on a single project as well (political science and business & management are DARPA SCORE alone, sociology is Student Replication alone, and education is 15 of 16 rows from the L2 Replication project), so they describe that project rather than the field. Table A is shorter still, because it additionally requires enough harvested rows coded direct: linguistics has only 10 of its 121, and the remaining fields fewer than 10. It is not restricted on the initiative side because several projects (SSRP, LOOPR, EROE, Sensory Marketing) record no rows coded direct, so restricting that column would delete those initiatives outright rather than make them comparable. "Biomedical" here means the database's biology discipline — cancer, cell, molecular biology, genetics and physiology — plus the 29 rows the database labels neuroscience that belong to the Brazilian Reproducibility Initiative. BRI is a single preclinical biomedicine project that stratified its sample on three common laboratory techniques — an MTT cell-viability assay, RT-PCR, and the elevated plus maze — and only that third arm carries a neuroscience label. Splitting one project across two disciplines by bench technique produced a spurious "neuroscience" row in earlier versions of these tables, in which 29 rodent anxiety experiments (14 distinct papers, each re-run in up to three laboratories) were compared against a harvested column that is mostly human fMRI and EEG work. Grouping the three arms together as biomedical also describes the harvested side of the row better than "biology" does, since it is largely human genetics and physiology from Nature Genetics, the American Journal of Human Genetics and the NEJM. Three fields are absent from both tables because only one of the two columns exists for them, so no comparison is possible: neuroscience (549 harvested rows, and — once BRI sits with the rest of its project — no initiative rows at all), medical fields (1,204 harvested rows, no initiative rows), and sports & exercise science (23 initiative rows from the Replicability of Sports and Exercise Science Research project, no harvested ones). The first two are the largest gaps in initiative coverage in the database. Education, political science and sociology appeared in earlier versions of this table but have been removed: their representative column consisted almost entirely of DARPA SCORE rows, and once SCORE's secondary-data reanalyses were removed from the database in August 2026 those cells fell to n = 16, 4 and 2 respectively. 3ie is excluded throughout: it re-analyses the original authors' existing datasets rather than collecting new data, which is reproducibility, not replication.
Which initiatives feed each representative cell (initiative → definitive n). Several projects span multiple fields, so their rows are assigned by each paper's home discipline rather than lumped under one label:
- Psychology (551): Student Replication Projects 145, Life Outcomes of Personality 118, RP:Psychology 96, DARPA SCORE 84, X-Phi 40, EROE 26, Social Science Replication Project 21, Sensory Marketing 21
- Biomedical (227): Cancer Biology 132, Brazilian Reproducibility Initiative 95 (all three of its technique arms)
- Linguistics (59): L2 Replication (Marsden et al.) 50, Student Replication 9
- Economics (29): Experimental Economics Replication Project 18, DARPA SCORE 11
- Education (16): L2 Replication (Marsden et al.) 15, DARPA SCORE 1
- Political science (4) and business & management (4): DARPA SCORE only · Sociology (2): Student Replication Projects only
Psychology has by far the most rows on both sides of both tables, and it is where the two views disagree — which is the point. Compared against everything we have harvested, psychology initiatives (58%) and harvested replications (59%) look identical. Compared against only our direct replications, the harvested rate drops to 49% and the initiatives come out ahead. The harvested corpus's apparent parity is therefore substantially a design effect: it is only 13% direct replications, and looser designs succeed more often (close extensions 66% vs direct 42%). The biomedical row shows the same movement — 50% on all types, 38% on direct only, against the initiatives' 31%.
Three cautions on reading these tables. First, the direct-only column is substantial only in psychology (n=302) and the biomedical row (n=82); economics (n=17) and education (n=21) sit at or below the threshold and should be read as blank, and linguistics is absent from Table A altogether because only 10 of its harvested rows are direct. Second, the two columns still differ in what was sampled, not just how it was replicated: the initiatives each enumerated a defined slice of literature (specific journals and years), while the harvested rows have no sampling frame at all, so this is a comparison across populations and not a clean bias estimate. Third, the initiative column counts replicate attempts, not distinct findings. Several projects re-ran the same target in multiple laboratories or recorded several outcomes per paper, so its 1,061 rows come from just 594 original papers — 1.79 rows per paper, against 1.36 in the harvested column, and 8.17 in Cancer Biology alone. The rates barely move when recomputed per paper (psychology 58% → 54%, biomedical 31% → 41%), but every n in the initiative column is larger than the number of independent findings behind it.
Note also that DARPA SCORE is deliberately cross-field, sampling 62 social-science journals, so its papers are split to their home disciplines above rather than lumped under one label.
How to read these numbers
- This is not a clean bias estimate. The two columns differ in what was sampled, not only in how it was replicated. A pre-specified frame over 62 social-science journals and an unframed harvest of whatever called itself a replication are different populations; the comparison bounds the problem rather than measuring it.
- The initiative buckets are a judgment call. "Representative" and "targeted" are assigned per project, and the assignment is contestable.
data/previous_replication_initiatives.csvrecords aselection_designfor each initiative; restricting the representative bucket to only its strict sampling frames (random_frame,exhaustive_frame,quasi_random_frame_feasibility,method_stratified) gives 43% (n=358) against the harvested 59% — the gap widens to 16 points rather than closing, so the direction of the finding survives the stricter definition. Reproduce this withpython3 scripts/build_bias_check.py --strict. - The unit is a replication attempt, not an independent finding. Coordinated projects contribute several rows per target — one per participating laboratory, or one per recorded outcome — so the initiative column's denominators are inflated relative to the harvested column's. Recomputing every rate at the level of the original paper leaves the picture intact (psychology 58% → 54%, biomedical 31% → 41%, economics 59% → 59%) but shrinks the initiative n's by roughly half, which is the honest measure of how much evidence stands behind them.
- Convenience corpus, non-random. Absolute rates describe this collection, not "the fraction of all science that replicates."
- Roughly two-thirds of rows are AI-curated and unvalidated, and "success" is heterogeneous — sometimes the replication authors' own judgment, sometimes a human curator's, sometimes the AI's. See Replication outcome classification.
- Discipline and initiative are confounded. Field-level rates partly reflect which projects sampled which fields, not intrinsic field differences.
These figures will be updated as the database grows and as more rows are human-validated. If you find an error or want to contribute data, see the replications database.