Correlates of Reproducibility

Each page of the replications database looks at one predictor of replication success at a time. This page puts them side by side: first as simple correlations between each predictor and the replication outcome, then jointly in an interpretable regression model that shows how much each factor contributes once the others are held fixed.

How the outcome is coded on this page

Analyses here score each replication attempt as 1 = success, 0 = failure or reversal, and 0.5 = recorded as inconclusive, using the stored result (the “reported” criterion; rows with no recorded result are dropped, n = 138). This differs from the site-wide success-rate definition (“success / (success + failure + reversal); inconclusive and unrecorded outcomes excluded”), which excludes inconclusive attempts entirely — see how outcomes are classified. Analyses are at the effect level; because one original paper can be replicated many times, all confidence intervals are paper-cluster bootstrap intervals (1,000 draws, resampling original papers), which are wider and more honest than intervals that assume independent rows.

Correlation table

Pearson r is computed on the transformed predictor shown in the second column (heavily skewed predictors are log-scaled); Spearman's ρ is rank-based, so it uses the raw values and is unaffected by any monotone transform. Each row links to the page that examines that predictor in depth, and each uses every effect for which that predictor is available, so n varies by row.

PredictorPearson scalen (effects)Papers95% CI95% CI
Original p-valuelog₁₀ p (floored at −10)656471-0.176[-0.274, -0.083]-0.130[-0.228, -0.034]
Original publication yearraw8,2665,7790.016[-0.025, 0.053]0.056[0.022, 0.088]
Journal impact factorlog₁₀ IF7,4865,155-0.080[-0.114, -0.046]-0.077[-0.112, -0.040]
Journal rank (SJR percentile)raw7,5935,228-0.082[-0.112, -0.048]-0.113[-0.144, -0.083]
Citation count (first 2 years)log₁₀(1+c)7,1634,861-0.084[-0.116, -0.049]-0.089[-0.122, -0.053]
First-author h-indexlog₁₀(1+h)7,1084,826-0.055[-0.092, -0.019]-0.051[-0.089, -0.011]
Last-author h-indexlog₁₀(1+h)7,0614,789-0.047[-0.091, 0.001]-0.054[-0.095, -0.011]
Mean author h-indexlog₁₀(1+h)7,1084,826-0.085[-0.117, -0.054]-0.083[-0.116, -0.050]
Max author h-indexlog₁₀(1+h)7,1084,826-0.075[-0.108, -0.045]-0.073[-0.111, -0.041]
Author overlap (shared authors)raw7,7715,4540.179[0.152, 0.209]0.280[0.254, 0.303]
  • Outcome coded 1 / 0.5 / 0 as described above; 95% CIs are paper-cluster bootstrap percentile intervals.
  • The p-value row uses only exactly reported p-values (type “=”, 656 effects) — stricter than the by-p-value page, which also admits p-values whose reporting type is unrecorded. Reported p-values run down to the smallest representable number, so for the Pearson correlation log₁₀ p is floored at -10 (p < 10-10 treated as 10-10); Spearman's ρ is unaffected.
  • SJR percentile is oriented so higher = better-ranked journal. Journals whose OpenAlex impact factor is exactly 0 are excluded from the impact-factor row (358 effects), since a log scale cannot represent them.

Summary table

The same estimates in a compact, shareable form.

Correlates of reproducibility

across 8,307 replicated effects from 5,813 papers

The Metascience Observatory
Predictorr95% CISpearman ρ95% CI
Author overlap (shared authors)0.18[0.15, 0.21]0.28[0.25, 0.30]
Original publication year0.02[-0.03, 0.05]0.06[0.02, 0.09]
First-author h-index-0.05[-0.09, -0.02]-0.05[-0.09, -0.01]
Last-author h-index-0.05[-0.09, 0.00]-0.05[-0.10, -0.01]
Max author h-index-0.08[-0.11, -0.04]-0.07[-0.11, -0.04]
Journal impact factor-0.08[-0.11, -0.05]-0.08[-0.11, -0.04]
Mean author h-index-0.09[-0.12, -0.05]-0.08[-0.12, -0.05]
Citation count (first 2 years)-0.08[-0.12, -0.05]-0.09[-0.12, -0.05]
Journal rank (SJR percentile)-0.08[-0.11, -0.05]-0.11[-0.14, -0.08]
Original p-value-0.18[-0.27, -0.08]-0.13[-0.23, -0.03]

What the correlations say

Every prestige-flavored predictor — journal impact factor, journal rank, citations, and author h-index in all four variants — correlates negatively with replication: more prestigious venues and more eminent, more-cited work replicate slightly less often. The single strongest correlate is author overlap (ρ = 0.280): replications that share authors with the original succeed far more often than independent ones. Smaller original p-values predict success (ρ = -0.130 on the exactly-reported subset), and publication year has at most a weak positive trend. All of these are small by conventional standards — no single feature of a paper comes close to determining whether it replicates.

A simple joint model

The correlations above are pairwise, and the predictors overlap (highly cited papers sit in high-impact journals written by high-h authors). To see what each factor contributes with the others held fixed, we fit a fractional logistic regression of the 0 / 0.5 / 1 outcome on all predictors at once. Covariates are standardized, so coefficients read as log-odds per one standard deviation and are directly comparable; the AME column converts each into percentage points of predicted replication score. Only mean author h-index enters (the four h-index variants are too collinear to separate, r ≈ 0.9), and the p-value — missing for most rows — is added only in Model B, which is restricted to the much smaller exactly-reported-p subset.

Model A — all effects with complete predictors

Publication year, journal impact factor and rank, first-2-year citations, mean author h-index, and author overlap.

Covariateβ per +1 SDSEzpAME (pp per +1 SD)+1 SD means…
Publication year0.0640.0391.630.1031.510.3 years later
log₁₀ impact factor0.0040.0450.090.9240.1≈ ×2.4 higher impact factor
SJR percentile-0.1050.043-2.430.015-2.510.5 SJR percentile points higher
log₁₀(1+cit. first 2 yrs)-0.0730.044-1.640.100-1.7≈ ×3.5 the first-2-year citations (+1)
log₁₀(1+mean author h)-0.0740.035-2.090.037-1.8≈ ×1.9 the mean author h-index (+1)
Shared authors0.4660.0617.63< 0.00111.11.8 more shared authors

Fractional logit (quasi-MLE), n = 6,305 effects in 4,293 original papers; standard errors clustered on the original paper. Covariates are z-scored on this model's own estimation sample, so coefficients are comparable within the table. AME = average marginal effect, in percentage points of predicted replication score per +1 SD.

Model B — adds the original p-value

Same covariates plus log₁₀ p (floored at -10), on the small subset with an exactly reported original p-value — read with caution.

Covariateβ per +1 SDSEzpAME (pp per +1 SD)+1 SD means…
Publication year0.0810.1270.640.5221.85.6 years later
log₁₀ impact factor0.0240.1420.170.8670.5≈ ×2.4 higher impact factor
SJR percentile0.1270.1231.030.3022.99.5 SJR percentile points higher
log₁₀(1+cit. first 2 yrs)-0.0540.139-0.390.698-1.2≈ ×3.4 the first-2-year citations (+1)
log₁₀(1+mean author h)-0.1610.131-1.230.219-3.6≈ ×1.9 the mean author h-index (+1)
Shared authors0.1920.1141.690.0904.32.8 more shared authors
log₁₀ p-value-0.4660.114-4.07< 0.001-10.5≈ ×440 larger p-value

Fractional logit (quasi-MLE), n = 451 effects in 357 original papers; standard errors clustered on the original paper. Covariates are z-scored on this model's own estimation sample, so coefficients are comparable within the table. AME = average marginal effect, in percentage points of predicted replication score per +1 SD.

Coefficients at a glance

Model A (n = 6,305)Model B (n = 451)-0.75-0.5-0.2500.250.5Standardized coefficient (log-odds per +1 SD) with 95% CI · right of 0 → more replicablePublication yearPublication year — Model A (n = 6,305): β = 0.064 per +1 SD [-0.013, 0.140], p = 0.103; AME 1.5 pp per +1 SDPublication year — Model B (n = 451): β = 0.081 per +1 SD [-0.167, 0.329], p = 0.522; AME 1.8 pp per +1 SDlog₁₀ impact factorlog₁₀ impact factor — Model A (n = 6,305): β = 0.004 per +1 SD [-0.084, 0.092], p = 0.924; AME 0.1 pp per +1 SDlog₁₀ impact factor — Model B (n = 451): β = 0.024 per +1 SD [-0.255, 0.303], p = 0.867; AME 0.5 pp per +1 SDSJR percentileSJR percentile — Model A (n = 6,305): β = -0.105 per +1 SD [-0.190, -0.020], p = 0.015; AME -2.5 pp per +1 SDSJR percentile — Model B (n = 451): β = 0.127 per +1 SD [-0.114, 0.369], p = 0.302; AME 2.9 pp per +1 SDlog₁₀(1+cit. first 2 yrs)log₁₀(1+cit. first 2 yrs) — Model A (n = 6,305): β = -0.073 per +1 SD [-0.159, 0.014], p = 0.100; AME -1.7 pp per +1 SDlog₁₀(1+cit. first 2 yrs) — Model B (n = 451): β = -0.054 per +1 SD [-0.326, 0.218], p = 0.698; AME -1.2 pp per +1 SDlog₁₀(1+mean author h)log₁₀(1+mean author h) — Model A (n = 6,305): β = -0.074 per +1 SD [-0.143, -0.005], p = 0.037; AME -1.8 pp per +1 SDlog₁₀(1+mean author h) — Model B (n = 451): β = -0.161 per +1 SD [-0.418, 0.096], p = 0.219; AME -3.6 pp per +1 SDShared authorsShared authors — Model A (n = 6,305): β = 0.466 per +1 SD [0.346, 0.585], p < 0.001; AME 11.1 pp per +1 SDShared authors — Model B (n = 451): β = 0.192 per +1 SD [-0.030, 0.415], p = 0.090; AME 4.3 pp per +1 SDlog₁₀ p-valuelog₁₀ p-value — Model B (n = 451): β = -0.466 per +1 SD [-0.690, -0.241], p < 0.001; AME -10.5 pp per +1 SDThe Metascience Observatory

Points are standardized coefficients; horizontal lines are cluster-robust 95% confidence intervals. Model A, Model B. Hover a point for details.

Reading the model

Author overlap dominates Model A: it is by far the largest coefficient, echoing the by-author-overlap page. Journal rank and mean author h-index retain small negative effects once everything else is held fixed, while impact factor and citations — strongly correlated with both — contribute little of their own. In Model B the original p-value is the strongest predictor in the model, and it absorbs much of what the prestige variables appeared to carry; but Model B rests on a small, unrepresentative subset (initiatives that report exact p-values), so treat it as suggestive. As always, these are observational associations, not causal effects.

Coverage

Source: replications_database_2026_08_01_120356.csv8,445 effect rows, of which 8,307 have a recorded result (5,813 original papers). Predictor coverage among those rows: publication year 8,266; journal impact factor 7,486 (after excluding 358 with IF = 0); journal rank 7,593; citations in the first two years 7,163; author h-index 7,108 (first author 7,108, last author 7,061); author overlap 7,771; exactly reported p-value 656. Impact factors and citation counts come from OpenAlex, journal ranks from SCImago, and author h-indexes from SciSciNet; all joins are by DOI or normalized journal name. Statistics on this page are verified against an independent Python implementation (scripts/check_logit.py).