Correlates of Reproducibility

Each page of the replications database looks at one predictor of replication success at a time. This page puts them side by side: first as simple correlations between each predictor and the replication outcome, then jointly in an interpretable regression model that shows how much each factor contributes once the others are held fixed.

How the outcome is coded on this page

Analyses here score each replication attempt as 1 = success, 0 = failure or reversal, and 0.5 = recorded as inconclusive, using the stored result (the “reported” criterion; rows with no recorded result are dropped, n = 138). This differs from the site-wide success-rate definition (“success / (success + failure + reversal); inconclusive and unrecorded outcomes excluded”), which excludes inconclusive attempts entirely — see how outcomes are classified. Analyses are at the effect level; because one original paper can be replicated many times, all confidence intervals are paper-cluster bootstrap intervals (1,000 draws, resampling original papers), which are wider and more honest than intervals that assume independent rows.

Correlation table

Pearson r is computed on the transformed predictor shown in the second column (heavily skewed predictors are log-scaled); Spearman's ρ is rank-based, so it uses the raw values and is unaffected by any monotone transform. Each row links to the page that examines that predictor in depth, and each uses every effect for which that predictor is available, so n varies by row.

PredictorPearson scalen (effects)Papers95% CI95% CI
Original p-valuelog₁₀ p (floored at −10)662476-0.170[-0.256, -0.083]-0.122[-0.205, -0.026]
Original publication yearraw8,7016,0730.008[-0.028, 0.049]0.044[0.010, 0.080]
Journal impact factorlog₁₀ IF7,8465,380-0.073[-0.105, -0.042]-0.071[-0.104, -0.038]
Journal rank (SJR percentile)raw7,9575,460-0.081[-0.111, -0.048]-0.108[-0.138, -0.077]
Citation count (first 2 years)log₁₀(1+c)7,2194,829-0.079[-0.113, -0.045]-0.085[-0.118, -0.050]
First-author h-indexlog₁₀(1+h)7,1694,800-0.053[-0.091, -0.018]-0.049[-0.088, -0.011]
Last-author h-indexlog₁₀(1+h)7,1234,764-0.048[-0.092, -0.001]-0.054[-0.094, -0.012]
Mean author h-indexlog₁₀(1+h)7,1694,800-0.082[-0.113, -0.049]-0.078[-0.109, -0.046]
Max author h-indexlog₁₀(1+h)7,1694,800-0.071[-0.100, -0.039]-0.069[-0.099, -0.036]
Author overlap (shared authors)raw7,7365,4180.180[0.148, 0.210]0.280[0.255, 0.305]
  • Outcome coded 1 / 0.5 / 0 as described above; 95% CIs are paper-cluster bootstrap percentile intervals.
  • The p-value row uses only exactly reported p-values (type “=”, 662 effects) — stricter than the by-p-value page, which also admits p-values whose reporting type is unrecorded. Reported p-values run down to the smallest representable number, so for the Pearson correlation log₁₀ p is floored at -10 (p < 10-10 treated as 10-10); Spearman's ρ is unaffected.
  • SJR percentile is oriented so higher = better-ranked journal. Journals whose OpenAlex impact factor is exactly 0 are excluded from the impact-factor row (374 effects), since a log scale cannot represent them.

Summary table

The same estimates in a compact, shareable form.

Correlates of reproducibility

across 8,743 replicated effects from 6,108 papers

The Metascience Observatory
Predictorr95% CISpearman ρ95% CI
Author overlap (shared authors)0.18[0.15, 0.21]0.28[0.25, 0.31]
Original publication year0.01[-0.03, 0.05]0.04[0.01, 0.08]
First-author h-index-0.05[-0.09, -0.02]-0.05[-0.09, -0.01]
Last-author h-index-0.05[-0.09, -0.00]-0.05[-0.09, -0.01]
Max author h-index-0.07[-0.10, -0.04]-0.07[-0.10, -0.04]
Journal impact factor-0.07[-0.10, -0.04]-0.07[-0.10, -0.04]
Mean author h-index-0.08[-0.11, -0.05]-0.08[-0.11, -0.05]
Citation count (first 2 years)-0.08[-0.11, -0.05]-0.09[-0.12, -0.05]
Journal rank (SJR percentile)-0.08[-0.11, -0.05]-0.11[-0.14, -0.08]
Original p-value-0.17[-0.26, -0.08]-0.12[-0.21, -0.03]

What the correlations say

Every prestige-flavored predictor — journal impact factor, journal rank, citations, and author h-index in all four variants — correlates negatively with replication: more prestigious venues and more eminent, more-cited work replicate slightly less often. The single strongest correlate is author overlap (ρ = 0.280): replications that share authors with the original succeed far more often than independent ones. Smaller original p-values predict success (ρ = -0.122 on the exactly-reported subset), and publication year has at most a weak positive trend. All of these are small by conventional standards — no single feature of a paper comes close to determining whether it replicates.

A simple joint model

The correlations above are pairwise, and the predictors overlap (highly cited papers sit in high-impact journals written by high-h authors). To see what each factor contributes with the others held fixed, we fit a fractional logistic regression of the 0 / 0.5 / 1 outcome on all predictors at once. Covariates are standardized, so coefficients read as log-odds per one standard deviation and are directly comparable; the AME column converts each into percentage points of predicted replication score. Only mean author h-index enters (the four h-index variants are too collinear to separate, r ≈ 0.9), and the p-value — missing for most rows — is added only in Model B, which is restricted to the much smaller exactly-reported-p subset.

Model A — all effects with complete predictors

Publication year, journal impact factor and rank, first-2-year citations, mean author h-index, and author overlap.

Covariateβ per +1 SDSEzpAME (pp per +1 SD)+1 SD means…
Publication year0.0650.0401.640.1011.510.3 years later
log₁₀ impact factor0.0040.0450.080.9320.1≈ ×2.4 higher impact factor
SJR percentile-0.1030.044-2.350.019-2.410.5 SJR percentile points higher
log₁₀(1+cit. first 2 yrs)-0.0720.045-1.610.107-1.7≈ ×3.5 the first-2-year citations (+1)
log₁₀(1+mean author h)-0.0770.036-2.170.030-1.8≈ ×1.9 the mean author h-index (+1)
Shared authors0.4660.0617.61< 0.00111.11.8 more shared authors

Fractional logit (quasi-MLE), n = 6,282 effects in 4,270 original papers; standard errors clustered on the original paper. Covariates are z-scored on this model's own estimation sample, so coefficients are comparable within the table. AME = average marginal effect, in percentage points of predicted replication score per +1 SD.

Model B — adds the original p-value

Same covariates plus log₁₀ p (floored at -10), on the small subset with an exactly reported original p-value — read with caution.

Covariateβ per +1 SDSEzpAME (pp per +1 SD)+1 SD means…
Publication year0.0890.1280.700.4842.05.6 years later
log₁₀ impact factor0.0310.1460.220.8290.7≈ ×2.4 higher impact factor
SJR percentile0.1260.1231.030.3042.99.5 SJR percentile points higher
log₁₀(1+cit. first 2 yrs)-0.0240.142-0.170.865-0.5≈ ×3.4 the first-2-year citations (+1)
log₁₀(1+mean author h)-0.1730.131-1.310.189-3.9≈ ×1.9 the mean author h-index (+1)
Shared authors0.1900.1141.660.0974.32.8 more shared authors
log₁₀ p-value-0.4380.114-3.84< 0.001-9.9≈ ×430 larger p-value

Fractional logit (quasi-MLE), n = 451 effects in 357 original papers; standard errors clustered on the original paper. Covariates are z-scored on this model's own estimation sample, so coefficients are comparable within the table. AME = average marginal effect, in percentage points of predicted replication score per +1 SD.

Coefficients at a glance

Model A (n = 6,282)Model B (n = 451)-0.5-0.2500.250.5Standardized coefficient (log-odds per +1 SD) with 95% CI · right of 0 → more replicablePublication yearPublication year — Model A (n = 6,282): β = 0.065 per +1 SD [-0.013, 0.142], p = 0.101; AME 1.5 pp per +1 SDPublication year — Model B (n = 451): β = 0.089 per +1 SD [-0.161, 0.340], p = 0.484; AME 2.0 pp per +1 SDlog₁₀ impact factorlog₁₀ impact factor — Model A (n = 6,282): β = 0.004 per +1 SD [-0.085, 0.092], p = 0.932; AME 0.1 pp per +1 SDlog₁₀ impact factor — Model B (n = 451): β = 0.031 per +1 SD [-0.255, 0.318], p = 0.829; AME 0.7 pp per +1 SDSJR percentileSJR percentile — Model A (n = 6,282): β = -0.103 per +1 SD [-0.188, -0.017], p = 0.019; AME -2.4 pp per +1 SDSJR percentile — Model B (n = 451): β = 0.126 per +1 SD [-0.114, 0.367], p = 0.304; AME 2.9 pp per +1 SDlog₁₀(1+cit. first 2 yrs)log₁₀(1+cit. first 2 yrs) — Model A (n = 6,282): β = -0.072 per +1 SD [-0.159, 0.016], p = 0.107; AME -1.7 pp per +1 SDlog₁₀(1+cit. first 2 yrs) — Model B (n = 451): β = -0.024 per +1 SD [-0.302, 0.254], p = 0.865; AME -0.5 pp per +1 SDlog₁₀(1+mean author h)log₁₀(1+mean author h) — Model A (n = 6,282): β = -0.077 per +1 SD [-0.147, -0.007], p = 0.030; AME -1.8 pp per +1 SDlog₁₀(1+mean author h) — Model B (n = 451): β = -0.173 per +1 SD [-0.430, 0.085], p = 0.189; AME -3.9 pp per +1 SDShared authorsShared authors — Model A (n = 6,282): β = 0.466 per +1 SD [0.346, 0.585], p < 0.001; AME 11.1 pp per +1 SDShared authors — Model B (n = 451): β = 0.190 per +1 SD [-0.034, 0.414], p = 0.097; AME 4.3 pp per +1 SDlog₁₀ p-valuelog₁₀ p-value — Model B (n = 451): β = -0.438 per +1 SD [-0.661, -0.215], p < 0.001; AME -9.9 pp per +1 SDThe Metascience Observatory

Points are standardized coefficients; horizontal lines are cluster-robust 95% confidence intervals. ● Model A, ◆ Model B. Hover a point for details.

Reading the model

Author overlap dominates Model A: it is by far the largest coefficient, echoing the by-author-overlap page. Journal rank and mean author h-index retain small negative effects once everything else is held fixed, while impact factor and citations — strongly correlated with both — contribute little of their own. In Model B the original p-value is the strongest predictor in the model, and it absorbs much of what the prestige variables appeared to carry; but Model B rests on a small, unrepresentative subset (initiatives that report exact p-values), so treat it as suggestive. As always, these are observational associations, not causal effects.

Coverage

Source: replications_database_2026_10_01_130503.csv — 8,881 effect rows, of which 8,743 have a recorded result (6,108 original papers). Predictor coverage among those rows: publication year 8,701; journal impact factor 7,846 (after excluding 374 with IF = 0); journal rank 7,957; citations in the first two years 7,219; author h-index 7,169 (first author 7,169, last author 7,123); author overlap 7,736; exactly reported p-value 662. Impact factors and citation counts come from OpenAlex, journal ranks from SCImago, and author h-indexes from SciSciNet; all joins are by DOI or normalized journal name. Statistics on this page are verified against an independent Python implementation (scripts/check_logit.py).