Published at MetaROR

August 10, 2026

Table of contents

Cite this article as:

Jones, L., Barnett, A., Hartel, G., & Vagenas, D. (2026). An Empirical Assessment of Inferential Reproducibility of Linear Regression in Health and Biomedical Research Papers. medRxiv, 2026-04.

An empirical assessment of inferential reproducibility of linear regression in health and biomedical research papers

Lee Jones123 EmailORCID, Adrian Barnett2 ORCID, Gunter Hartel3 ORCID, Dimitrios Vagenas1

1 Research Methods Group, Faculty of Health, School of Public Health and Social Work, Queensland University of Technology, Kelvin Grove, Queensland, Australia,
2 AusHSI, Centre for Healthcare Transformation, Faculty of Health, School of Public Health and Social Work, Queensland University of Technology, Kelvin Grove, Queensland, Australia,
3 Statistics Unit, QIMR Berghofer Medical Research Institute, Herston, Queensland, Australia,

Originally published on April 7, 2026 at: 

Abstract

Background: In health research, variability in modelling decisions can lead to different conclusions even when the same data are analysed, a challenge known as inferential reproducibility. In linear regression analyses, incorrect handling of key assumptions, such as normality of the residuals and linearity, can undermine reproducibility. This study examines how violations of these assumptions influence inferential conclusions when the same data are reanalysed.

Methods: We randomly sampled 95 health-related PLOS ONE papers from 2019 that reported linear regression in their methods. Data were available for 43 papers, and 20 were assessed for computational reproducibility, with three models per paper evaluated. The 14 papers that included a model at least partially computationally reproduced were then examined for inferential reproducibility. To assess the impact of assumption violations, differences in coefficients, 95% confidence intervals, and model fit were compared.

Results: Of the fourteen papers assessed, only three were inferentially reproducible. The most frequently violated assumptions were normality and independence, each occurring in eight papers. Violations of independence were particularly consequential and were commonly associated with inferential failure. Although reproduced analyses often retained the same binary statistical significance classification as the original studies, confidence intervals were frequently wider, indicating greater uncertainty and reduced precision. Such uncertainty may affect the interpretation of results and, in turn, influence treatment decisions and clinical practice.

Conclusion: Our findings demonstrate that substantial violations of key modelling assumptions often went undetected by authors and peer reviewers and, in many cases, were associated with inferential reproducibility failure. This highlights the need for stronger statistical education and greater transparency in modelling decisions. Rather than applying rigid or misinformed rules, such as incorrectly testing the normality of the outcome variable, researchers should adopt modelling frameworks guided by the research question and the study design. When assumptions are violated, appropriate alternatives, such as robust methods, bootstrapping, generalized linear models, or mixed-effects models, should be considered. Given that assumption violations were common even in relatively simple regression models, early and sustained collaboration with statisticians is critical for supporting robust, defensible, and clinically meaningful conclusions.

Introduction

The reproducibility of health research is essential for assessing the reliability of findings and supporting informed clinical decision-making. When research is reproducible, healthcare professionals can be more confident that their decisions are based on trustworthy evidence, thereby improving outcomes and reducing the risk of harm from ineffective or inappropriate interventions [1]. However, many published health and biomedical studies cannot be reliably reproduced, contributing to what is now widely recognised as a reproducibility crisis [2, 3].

Reproducibility concerns highlight the need to strengthen research standards through rigorous study design, appropriate statistical methods, and transparent reporting. Empirical assessments by Hardwicke et al. [4] and others have shown that essential materials, including data, code, and analysis plans, are frequently unavailable or insufficiently documented, limiting the ability to reproduce published findings [5]. These challenges raise concerns about the reliability of scientific evidence and emphasise the importance of reproducibility in supporting sound healthcare decisions and safe, effective patient care, enabling a more efficient use of limited resources on interventions more likely to benefit public health [6].

In this paper, we focus on inferential reproducibility, defined as the consistency of conclusions drawn from the same dataset when alternative, reasonable analytical choices are applied [3]. Although methodological and computational reproducibility have received increasing attention, inferential reproducibility remains comparatively underexplored [3]. Differences in analytical decisions, such as the choice of statistical test, covariate adjustment, model specification, or outcome prioritisation, may be legitimate in principle. However, when these decisions lack clear justification or transparency, they can lead to divergent interpretations of the same data [7, 8], undermining confidence in research findings and complicating the identification of effective treatments.

The reproducibility crisis is further exacerbated by limited transparency in how statistical conclusions are derived. Researchers may fail to document their analytical rationale or consider alternative approaches when assumptions are violated, and formal assessment of assumptions is often absent from publications [9]. Inconsistent statistical practices, gaps in training, and limited access to methodological expertise further contribute to variability in statistical inferences and conclusions. Publication incentives that prioritise novelty and statistical significance may also encourage selective reporting and over-interpretation [10].

These challenges are particularly relevant for linear regression, one of the most widely used statistical methods in health research [11]. The validity of linear regression relies on assumptions of linearity, independence, homoscedasticity, and normality of residuals. Violations of these assumptions can meaningfully alter effect estimates, standard errors, and statistical inference [12, 13]. This study examines how such violations contribute to variability in inferential reproducibility, in which analytical decisions can lead to differing conclusions despite using identical data. It provides a practical guide for assessing statistical assumptions and offers recommendations to improve transparency, promote rigour, and support more reliable evidenced based decision making.

Materials and Methods

This study is the second part of a project on statistical quality in published research. In Stage 1, experienced statisticians reviewed selected papers for linear regression reporting practices, examining whether authors reported checking statistical assumptions and how those assumptions were assessed [14, 15]. Stage 2 examines reproducibility. The first component of Stage 2 assessed computational reproducibility, examining whether data were available and whether published results could be reproduced from the available materials [16]. This paper presents the second component of Stage 2, assessing inferential reproducibility, specifically, whether the reported conclusions are justified based on the statistical assumptions. This study received negligible-low risk ethics approval from the Queensland University of Technology Human Research Ethics Committee (Approval No. 2000000458).

Sample size determination

A sample of 100 papers was selected to evaluate the prevalence of reported linear regression assumptions. The sample size was calculated to estimate a proportion of 0.05 (5%) with a two-sided 95% confidence interval and a margin of error of 5% [14]. It was anticipated that approximately 50% of papers would contain accessible data, enabling reproducibility assessment. This study was designed as a pilot, given the current lack of empirical evidence on the proportion of published studies that are reproducible.

Selection of papers

The initial sample of 100 research articles, excluding editorials and other non-research content, was randomly selected from PLOS ONE publications in 2019 to provide a relevant snapshot of reporting practices. This time-frame also offered an unintended advantage: it was the last full year prior to the COVID-19 pandemic, thereby avoiding potential disruptions to research [17].

The selection used the ‘searchplos’ function from the ‘rplos’ R package [18] to identify articles that met two criteria: (1) the subject area included “health,” and (2) the Materials and Methods section reported the use of “linear regression.” Eligible papers were randomly ordered, and the first 100 meeting the inclusion criteria were selected. To focus on standard applications of linear regression in health research, articles were excluded if they used clustered or random effects, alternative modelling approaches (e.g., Bayesian or non-parametric methods), or applied linear regression only for secondary purposes (e.g., calibration or pre-processing). Upon full-text review, five papers were excluded as they did not report any linear regression results, leaving 95 papers for the assessment of reporting practices.

Research questions

  • What proportion of analyses and papers violated statistical assumptions (independence, linearity, homoscedasticity, and normality of residuals), including issues with outliers or multicollinearity?
  • How many papers reported that the authors had checked statistical assumptions, outliers, and multicollinearity?
  • What proportion of analyses remained inferentially reproducible after accounting for unmet assumptions, outliers, and multicollinearity?

    Evaluating Model Assumptions and Reproducibility

    Reproducibility analyses were conducted by the primary author, an accredited senior biostatistician with over 20 years of experience in statistical modelling and analysis of healthcare data. To ensure the reliability of the findings, selected results were independently reviewed by three senior biostatisticians, two of whom are accredited (AStat). Agreement on the reproduced results and interpretation was reached through discussion and consensus. The analyses followed a consistent, transparent process, documenting all steps and decisions. We refer to each paper by a unique number, the full list of identifiers and corresponding DOIs are available in the GitHub data [19].

    A maximum of three linear regression analyses per paper was selected to ensure the workload remained feasible. Because models within a paper typically share the same data and analytical framework, their reproducibility was expected to be highly correlated, limiting the additional insight gained from reproducing all models. For papers with more than three models, the final or primary models of interest were included first, with any remaining models selected at random. If no primary model was identified, all three were selected at random. Models that were partially reproducible were further assessed for inferential reproducibility, using the reproduced results and available data to check model assumptions and conduct sensitivity analyses.

    The assumptions of linear regression were evaluated using descriptive statistics, diagnostic plots, and formal tests, with additional checks for multicollinearity, outliers, and influential observations. A simulated example illustrating the relationship between blood pressure and age is shown in Fig 1 to demonstrate how these diagnostic procedures were applied in practice. While full implementation details are provided below and on GitHub [19], this example highlights key components of assumption checking, including inspection of scatter plots to assess linearity and potential non-linear (e.g., quadratic) relationships, examination of residual plots for evidence of non-linearity or heteroscedasticity, identification of influential observations, and assessment of whether residuals are approximately normally distributed.

    Figure 1. Simulated example showing a hypothetical relationship between blood pressure and age. The top-left scatter plot shows a clear positive relationship between age and blood pressure, with the blue line showing a linear fit and the red line a GAM fit. The residual plot (top-right) indicates a quadratic pattern that is not captured by the linear model but does not show evidence of heteroscedasticity. The Cook’s distance plot (bottom-left) shows that observations 139 and 191 exert greater influence on the model, though the overall Cook’s distances remain low, suggesting that outliers are not a major concern. The Q–Q plot (bottom-right) indicates that the residuals from the linear fit are approximately normally distributed.

    Visualisations of the model were generated using the ‘visreg’ package in R [20]. For linear predictors, each plot displays the relationship between the predictor (x-axis) and the outcome (y-axis), with other covariates held constant. Results for each paper are available on GitHub [19]. Each regression analysis was assessed using the steps outlined below:

    Is the linear structure of the regression model appropriately specified?

    • Scatterplots of standardized residuals versus predicted values and versus each predictor were used to assess the linearity of the overall model. Quadratic fits were overlaid to detect potential non-linear patterns [21].
    • Residual specification tests for predictors were conducted using quadratic terms. For the overall fitted values, Tukey’s one-degree-of-freedom test was used to assess curvature and non-additivity [21].
    • Scatterplots of raw dependent and predictor variables were examined with overlaid linear fits and smoothed generalized additive model (GAM) fits to explore potential non-linear relationships [22].

      Do residuals display homogeneity of variance?

      • The plot of residuals vs predicted values was assessed for patterns such as funnelling.
      • Plots of residuals vs individual variables were used to assess patterns.
      • Heteroscedasticity was assessed using the studentized Breusch–Pagan test for non constant error variance [23].

        Were there outliers and influential observations?

        • Cook’s distance measured the overall change in fit if the ith observation is removed. Potential influential observations are identified by
        • Cook’s Embedded Image, where n is the number of observations [24]. A threshold of 0.5 was used to identify influential observations.
        • DFFITS were used to assess how many standard deviations the fitted value for an observation changes when that observation is removed from the model [24]. A practical cut-off of 1 was used to identify observations with substantial influence.
        • DFBETAS were used to assess how many standard errors a regression coefficient changes when an observation is removed from the model [24]. A practical cut-off of 1 was used to flag observations with a meaningful impact.
        • Influence plots were examined to identify observations with both high leverage and large studentized residuals (typically ±2 or ±3), as these may disproportionately affect the model’s estimates. Such points appear as large, dark blue bubbles in the plots [21].
        • COVRATIO measures the overall change in the covariance structure of the estimated regression coefficients, reflecting changes in their precision when each observation is removed [25]. Values close to 1 indicate little influence on the model’s precision. Values below 1 indicate inflated variances and reduced precision (wider confidence intervals), whereas values above 1 indicate deflated variances and increased precision (narrower confidence intervals). A practical cut-off between 0.9 and 1.1 was applied, with values outside this range were considered to have a meaningful impact on precision.

          Are the residuals of the regression model approximately normally distributed?

          • A Q–Q plot of model residuals with confidence bands assessed how residuals compared to the normal distribution. Bands indicated expected variability, and points outside the bands suggested systematic deviations [21].
          • Residuals were described using descriptive statistics, including mean, standard deviation, median, minimum, maximum, skewness, and kurtosis.
          • Normality was assessed using the Shapiro–Wilk and Kolmogorov–Smirnov tests [26], alongside graphical and descriptive diagnostics. Recognising that such tests are highly sensitive in large samples and may detect minor deviations, while lacking power in small samples, statistical significance alone was not taken as evidence of a meaningful violation. While formal tests contributed to the overall assessment, they were also examined to understand how often they reject minor departures from normality.

            Are the residuals of the model independent of each other?

            • Each paper’s study design was reviewed to assess whether independence was theoretically plausible, with particular attention to potential clustering e.g., repeated measures within individuals or observations nested within hospitals.
            • Serial correlation was assessed using the Durbin–Watson test [27].

              Is there collinearity among the independent variables?

              • Variance Inflation Factors (VIFs) were examined to assess multicollinearity [27], with values greater than five considered potentially problematic [28].
              • Changes or inflation in standard errors across regression models were monitored as indicators of multicollinearity or model instability.

                Inferential reproducibility

                Inferential reproducibility was assessed only for models that demonstrated sufficient computational agreement with the original results, defined as numerical agreement within pre-specified tolerance thresholds based on reported decimal precision [16]. Models classified as “Reproduced”, “Mostly reproduced”, or “Partially reproduced” were included, whereas models classified as “Not reproduced” were excluded, as meaningful evaluation of modelling assumptions requires adequate correspondence between the reproduced and original analyses. Models categorised as Partially reproduced were retained because they generally reflected the core structure of the original model, allowing assessment of whether violations of assumptions could plausibly have influenced the study conclusions. Inferential reproducibility was then evaluated as a binary outcome (Yes/No), based on whether alternative modelling approaches or assumption checks materially altered the study conclusions.

                Results were presented in structured HTML reports with four tabs: Original Results, Reproduced Results, Differences, and Sensitivity Analysis, allowing readers to navigate easily and compare findings. Original results were extracted from the text and tables of the published articles, with any unreported statistics left blank to highlight gaps. The Reproduced Results tab displayed re-estimated coefficients and model fit statistics (e.g., R2), along with a summary of model assumption checks. The Differences tab was based on agreement between original and reproduced results (e.g., regression coefficients, standard errors, confidence intervals, test statistics) which were assessed using tolerance thresholds based on reported decimal place precision, classifying values as “Reproduced”, “Incorrect rounding”, or “Not reproduced”. For computational reproducibility results, see Jones et al. [19].

                The Sensitivity Analysis tab of the HTML report presents inferential reproducibility assessments. Continuous dependent and independent variables from the reproduced models were standardized by centering and scaling to enable scale-free comparison of effect sizes. Standardized coefficients represent the expected change in the outcome, expressed in standard deviation units, for a one standard deviation increase in the predictor. Categorical variables were not standardized and were left in their original form (dummy variables). Standardization allowed coefficients to be compared on the same scale; however, comparing coefficients was not always an appropriate indicator of non-reproducibility, particularly when key model assumptions, such as linearity or distributional assumptions, were violated.

                Differences in standardized regression coefficients were evaluated by comparing the reproduced model with sensitivity analyses assessing model assumptions. Three thresholds were applied: a percentage change of 10%, and absolute differences of 0.1 (moderate) and 0.2 (substantial). The 10% threshold was included as it was specified in the study protocol [29] and has been used in several previous studies [30, 31]. However, the percentage change can be misleading when regression coefficients are close to zero, as small absolute differences can appear disproportionately large. Because this was an exploratory project, we expanded on our protocol by also including absolute differences of 0.1 and 0.2 as more meaningful standardized measures. A difference of 0.1 can be considered similar in magnitude to a 10% change for many coefficients, while 0.2 aligns conceptually with Cohen’s d, where values below 0.2 are typically regarded as small [32]. Both percentage and absolute comparisons were explored to determine which thresholds more effectively captured meaningful differences and potential violations of assumptions.

                There were five general steps followed to assess inferential reproducibility:

                1. Independence design/cluster effects: If analyses were suspected of violating independence and an identification or clustering variable was available, the reproduced linear model (LM) and linear mixed model (LMM) were fit to compare the fixed effects [33]. The intraclass correlation coefficient (ICC) from the null LMM was reported [34]. Maximum-likelihood AIC values were then compared; when the LMM had a lower or similar AIC, it was retained as the primary model to preserve the data’s correlation structure for coefficient comparisons.

                2. Non-linearity: Linearity was assessed by adding quadratic terms (or other non-linear terms as required). Non-linear terms were retained only when they substantially improved model fit, as indicated by reductions in AIC and BIC values and by visual inspection that showed a better representation of the observed relationship.

                3. Bootstrap comparisons: Confidence intervals (95% CIs) for linear models were estimated using the bias-corrected and accelerated (BCa) bootstrap method [35]. Where BCa intervals failed to converge, percentile bootstrap intervals were used. For outcomes with non-independent observations, linear mixed models (LMMs) were fitted [33], and CIs were obtained using a parametric bootstrap based on model simulations (10,000 resamples). When heteroscedasticity was suspected, a wild bootstrap was applied in both LM and LMM settings. Bootstrap distributions were inspected for bimodality and heavy tails; where results appeared sensitive to outliers, targeted robust or sensitivity analyses were conducted.

                4. Comparison of results across models: Percentage and absolute changes in estimates and confidence-interval bounds relative to the reproduced linear model were summarised using thresholds of 10% change and standardized coefficient differences of <0.10 and <0.20. Coefficient direction and statistical significance were assessed for consistency.

                5. Distribution check: If inspection of model residuals or mean–variance relationships suggested that the assumed error distribution was inappropriate, sensitivity analyses were conducted using alternative distributions (e.g., Gamma for positively skewed continuous outcomes, Poisson or Negative Binomial for count data), as appropriate. Model assumptions, including the influence of outliers, were reassessed following refitting [36].

                To address assumption violations, we prioritised simple, consistent approaches that enabled meaningful comparisons across models. Bootstrapping was the primary method used to account for violations of the normality assumption, with wild bootstrap [37] applied when heteroscedasticity was suspected. When linearity was questionable, models were refitted with polynomial terms, and model fit statistics such as R2 and AIC/BIC, were compared rather than coefficient estimates. In cases where residuals were non-normal or showed a clear mean–variance relationship, models based on alternative theoretical distributions (e.g., Poisson, Gamma) were fitted [22]. Model fit was compared using the AIC and BIC. When multicollinearity was detected, regularisation methods such as ridge regression were applied to obtain more stable and reliable coefficient estimates [38].

                When studies reported removing outliers, we did not re-include them. This decision was based on two considerations: documentation surrounding outlier removal was often poor, and data for the outlier may not have been included in the dataset. To ensure valid comparisons of model fit, models were reproduced using the same dataset used for the original analysis. Recommendations regarding outlier handling were noted when relevant. Where necessary, alternative models were compared using information criteria, interpreted according to standard guidelines. For AIC, a reduction of less than 2 indicates support of a similar model, a reduction of 4–7 indicates improved fit, and a reduction greater than 10 indicates substantially better fit [39]. For BIC, Δ BIC values of 0–2 indicate weak evidence, 2–6 positive evidence, 6–10 strong evidence, and values greater than 10 very strong evidence in favour of the model with the lower BIC [40]. These thresholds are intended as general guidelines and were interpreted in the context of model purpose, plausibility, and diagnostic assessment.

                Due to our limited contextual knowledge of the original studies and our decision not to contact the authors, we retained the variables reported in each model, even when the models appeared to be overfit. In such cases, a recommendation to reduce the number of variables was noted. Models were classified as not inferentially reproducible when insufficient detail or key variables related to the study design were omitted, such as variables involving clustering, randomisation, or interactions. While efforts were made to standardize the overall approach, statistical analysis is inherently context-dependent. No single method is appropriate in all situations, and care was taken to select the most suitable approach for each case.

                Statistical methods

                The purpose of this study was descriptive rather than hypothesis test-driven. Descriptive statistics, including frequencies and percentages, were reported. Model-level classification of inferential reproducibility (Yes/No) was based on defined decision criteria. A standardized difference of 0.1 was used as the primary quantitative threshold, while a 10% change in key estimates and a standardized difference of 0.2 were reported descriptively for comparison. These quantitative criteria were considered alongside assessment of model assumptions (e.g., linearity and independence); the full decision framework is described above in the Inferential reproducibility section.

                At the model level, up to three models were assessed for inferential reproducibility per paper, resulting in clustering of models within papers so independence could not be assumed. Accordingly, a Bayesian mixed-effects logistic regression model for the binary outcome of inferential reproducibility was fitted, with a random intercept for paper, to estimate predicted probabilities. Weakly informative priors were used to stabilise estimation given the small number of papers [41]. The intraclass correlation coefficient (ICC) and its 95% credible interval were derived from the posterior distribution of the paper-level random effect using the standard latent-scale formulation for logistic mixed models [42]. Model adequacy and convergence were assessed using standard Bayesian diagnostics, including rank-normalised convergence statistics and trace plots [43]. At the paper level, inferential reproducibility was summarised as the proportion of papers in which all assessed models met the reproducibility criterion, with 95% Wilson confidence intervals used to quantify uncertainty around this estimate. R version 4.4.2 was used for all analyses [26].

                Results

                Of the 95 papers reviewed, 80 (84%) were observational studies, and 73 (77%) involved human participants. Among these papers, 68 reported having data available; however, the raw data necessary to reproduce the analyses could not be identified in 25 papers [16]. From the original random sample, a subset of 20 papers were sequentially assessed for computational reproducibility, of which eight were successfully reproduced, with a further six papers containing models that were at least partially reproducible.

                Therefore, 14 papers were evaluated for inferential reproducibility. Of these, three or 21% (95% CI: 8%, 48%) met the inferential reproducibility criterion, defined as all assessed models within the paper being reproducible (Fig 2).

                Figure 2. Flow chart of the included papers in the analysis. N is the number of papers.

                In total, 32 models across 14 papers were assessed for inferential reproducibility. Of these models, 14/32 (44%) were inferentially reproducible based on the observed data. After accounting for clustering by paper, the model-based marginal probability of reproducibility was 0.34 (95% Credible Intervals 0.07, 0.69). The reduction relative to the observed proportion reflects substantial between-paper variation, with inferential reproducibility tending to cluster within papers (ICC = 0.59, 95% CI: 0.05, 0.89). Given the small number of papers (N = 14) and the resulting uncertainty reflected in the wide credible interval, the estimate should be interpreted cautiously. Model diagnostics indicated satisfactory convergence. Posterior predictive checks suggested that the model adequately represented the observed data, see [

                Inferential reproducibility was assessed primarily by examining whether regression coefficients and their 95% confidence intervals fell within specified thresholds. This approach recognises that departures from model assumptions may arise from multiple interacting sources; therefore, specific violations cannot always be directly linked to discrepancies in estimates unless the connection is clear. Models with minor assumption departures were still considered inferentially reproducible when diagnostic assessments indicated these were unlikely to materially affect parameter estimates, interpretation, or overall model adequacy. Most authors reported limited information on statistical assumptions. Table 1 summarises the number of papers in which assumptions were checked or discussed, and the number in which at least one violation was identified in the reproduced analyses. Sample sizes differ between the reported and reproduced analyses: the reported results reflect the full paper, whereas the reproduced analyses were restricted to the three selected models per paper.

                 

                Table 1. Statistical assumption checks and violations across papers (N = 14).
                Paper reported: Authors reported checking or discussing the assumption; Reproduced analyses: Assumption violated in at least one of our reproduced analyses.
                Paper reported Reproduced analyses
                Assumptions N Checked N Violated
                Normality 14 8 (57%) 14 8 (57%)
                Linearity 12 2 (17%) 11 5 (45%)
                Homoscedasticity 14 3 (21%) 13 7 (54%)
                Independence 14 2 (14%) 14 8 (57%)
                Collinearity 11 4 (36%) 8 0 (0%)
                Outlier 14 6 (43%) 14 4 (29%)
                Residual normality was assessed using the Kolmogorov–Smirnov (KS) and Shapiro–Wilk (SW) tests, with Q-Q plots informing the final assessment across 32 models. In five models, all tests and plots indicated there was no violation of normality; in seven models, all indicated there were violations of normality. In seven models, one of the tests indicated a violation and the Q–Q plot also suggested non-normality. There was 13 cases where one of the tests indicated a violation, but the deviation on the Q-Q plot was negligible, and the residuals were judged to be approximately normal.

                The most frequently violated assumptions were residual normality and independence, each of which was violated in eight papers. Mild departures from normality did not always lead to inferential failure; however, violations of independence almost always did, with 16/32 models exhibiting independence violations. Gross violations of normality often suggested that an alternative outcome distribution would be more appropriate. The Gaussian (normal) distribution provided an adequate fit for most models (28) whereas a Tweedie distribution provided a better fit for three models and an Inverse Gaussian model for one. Linearity violations were identified in five papers, affecting 8 models. Further evaluation of model fit (AIC and BIC) and visual inspection of the data indicated that five of these models exhibited clear non-linearity. For descriptions and reproducibility results for individual papers, see Table 2.

                Table 2. Summary and paper-level reproducibility results
                Paper Summary and computational reproducibility results Inferential reproducibility results
                Paper 8 This study examined whether childhood emotional, physical, and sexual abuse were independently associated with depressive symptoms, emotion dysregulation, and interpersonal problems using linear regression. The authors reported checking nor­mality, linearity, homoscedasticity, and collinearity, but did not discuss outliers or provide details of these checks. The models assessed in this paper were mostly computationally reproducible, with minor errors identified. Some estimates had incorrect rounding, and in the emotional dysregulation model, the variable childhood sexual abuse had a lower 95% confidence interval that was missing a negative sign. The authors wrongly reported R2 rather than adjusted R2. The analyses were inferentially reproducible. The main mod­elling issue related to the distribution of the independent vari­ables: although predictor normality is not assumed in linear regression, sparse upper-end values and concentration at mini­mum scores affected leverage and linearity assessments. Mild violations of model assumptions were observed in the resid­uals. In the reproduced analyses, residual plots for all three models indicated mild heteroscedasticity, and Model 1 also showed mild non-normality. The third model exhibited po­tential departures from linearity and influential observations. However, refitting this model to include quadratic and cubic terms did not result in a meaningful improvement in model fit, as indicated by AIC, BIC, or adjusted R2. These issues were therefore likely attributable to sparse data at the upper end of the predictor scales. All three models were bootstrapped and produced results consistent with the reproduced models, with differences in standardized regression coefficients of less than 0.1. The direction and statistical significance of the regression coefficients remained consistent across both the reproduced and bootstrapped analyses.
                Paper 16 This study assessed the reliability of six activity monitors in patients with chronic heart failure relative to the Actigraph. The three devices chosen randomly for reproducibility were: Omron (OMR), SmartLAB walk+ (SLW), and Withings Go (WGO). The authors reported testing all data for normality using the Shapiro–Wilk test and removing an extreme outlier deemed a likely error; however, no other assumption checks were reported. The authors noted that the Bland–Altman plots revealed no systematic differences between the activity monitors and Actigraph. They presented the mean difference between devices and plots. For OMR and SLW models, the mean difference between the devices and the Actigraph (intercept model) was successfully reproduced; however, the p-values, which the authors implied were not significant, could not be
                replicated. The WGO model was not computationally reproducible, with inconsistent data between plots and tables. In contrast, the OMR and SLW models were partially reproduced.
                The devices were tested on participants over three days, so the measurements were not independent; therefore, linear mixed-effects models were used for sensitivity analyses for the OMR and SLW models. While the authors interpreted the plots, they reported Bland–Altman tests of the mean difference only in tables. These models were therefore the focus of the assessment of inferential reproducibility. The mean difference can also
                be tested using a paired t-test, which is equivalent to a one-sample t-test on the differences, and in a regression framework corresponds to a model with only an intercept. Standardized confidence interval widths differed (>0.2 for OMR, 0.15 for
                SLW), reflecting artificially narrow CIs when repeated measures were ignored in the authors model, as supported by high intraclass correlation coefficients (ICCs). To further assess bias, a second linear mixed-effects model was fitted to examine pro-
                portional bias. After adjusting for the mean of the two devices, there was no evidence of fixed bias; however, proportional bias was present, suggesting that differences between devices varied with the magnitude of measurement. OMR and SLW models were not considered inferentially reproducible because they did not account for repeated measures, with intraclass correlations of 0.35 and 0.26, respectively.
                Paper 19 A study examining associations between industry involvement and study characteristics of trial registration in biomedical research. Linear regression was a minor part of the paper, focusing on the mean difference for target sample size and industry involvement. The results for this paper were computationally reproduced. Results were presented in the text (co-efficient and 95% confidence interval), with additional decimal places provided in the supporting information for reproducibility checks. Reproduction was moderately difficult because the raw data required substantial manipulation and merging of four datasheets in both long and wide formats. A complete understanding also required consulting both the paper and the supporting information, as six outliers were excluded but not mentioned in the main text. The original analysis found a smaller target sample size for industry involvement (mean difference = -153, 95% CI -231 to -74, p < 0.001). A typographic error was identified in the supplementary results, where the upper confidence interval was reported as –74.98; when reproduced, it was -73.987, consistent with the correctly reported value (-74) in the main text. This model was found
                to be mostly reproducible.
                The regression in this paper was not inferentially reproducible, with gross violations of normality and heteroscedasticity. Removing outliers did not address the skewed distribution of target sample size, as indicated by large discrepancies between means and medians and SDs greatly exceeding IQRs (e.g., no industry: mean = 249, 95% CI 211–295, SD = 689; median = 70, IQR = 125; industry: mean = 96, 95% CI 26–166, SD = 198; median = 45, IQR = 76). Outlier removal reduced variance but risked bias, as large sample sizes are often required when rare events occur, or small effects are of interest. The target sample size was standardized, and a heteroscedastic bootstrap model was fitted, which showed a standardized change in confidence interval range of –0.1. While bootstrapping provided a comparison on the original additive scale, the clear mean–variance relationship (variance increasing with mean) indicated multiplicative rather than additive error. Alternative distributions were therefore considered. The negative binomial and gamma models improved fit, but residuals continued to show overdispersion. The inverse Gaussian model provided the best fit, with an AIC of 4,784 units lower than that of linear regression
                Paper 22 There were eight multivariable models identified in the paper, with the primary focus on physical, cognitive and emotional symptoms after traumatic brain injury (TBI). Three linear regressions were randomly sampled, including outcomes from the Rivermead Post-concussion Questionnaire (RPQ), both the full scale and the RPQ13 subscale, and the Impact of Event Scale (IES). All analyses were adjusted for age, ethnicity, and
                extracranial trauma. Assumptions, outliers, and multicollinearity were not discussed. Although the reproduced sample sizes matched those in the publication, minor discrepancies in medians and regression results suggested minor differences in the
                data. These may have arisen from updates to the dataset after results were generated. For instance, the median age for participants experiencing assault was reported as 29.3 (22.1, 36.4), whereas the reproduced median was 29.0 (22.0, 35.0). Linear regression coefficients and p-values often differed at the decimal level, with at least one aspect of each model being reproduced. The three models assessed in this paper were not
                computationally reproducible and were only partially reproduced.
                None of the three models assessed in this paper were inferentially reproducible. The primary reasons were violations of model assumptions and issues with data distribution and parameterisation. Ethnicity was treated as a continuous variable, even though it is nominal. Across the three outcome variables, 10–24% of values were zero. This was most pronounced in the IES model, where a strong residual pattern indicated violation of the normality assumption. A Tweedie distribution, which accommodates zero values and non-normality, provided a substantially better fit, with AIC reductions of 69–159 across outcomes, for RPQ_full and RPQ13. Models including a squared age term were explored to assess potential non-linearity. While model fit indices (AIC and BIC) showed marginal improvement, the differences were minor, and the Tweedie model with linear age was retained.
                Paper 23 This paper examines faecal contamination in various environmental samples in Bangladesh. The authors report using generalised linear models but do not specify the family or link function. In Stata, the Gaussian distribution is the default, and the data were log10-transformed before modelling, so the analysis was treated as a linear regression. The study design should have addressed the potential correlation of observations due to geographic clustering and sampling strategy. The tables do not indicate whether analyses were univariate or multivariable, and no continuous variables were included, so linearity was not required. No assumptions or outliers were discussed.
                The analysis in this paper was computationally reproducible, with differences from rounding errors. There were at least 35 regression models; the three randomly selected were univariate models predicting E. coli levels from different sources: drain water, water from produce, and drinking water. Models one and two compared neighbourhood type, and model three compared different sources of drinking water.
                This paper was not inferentially reproducible because the study was geographically clustered. Linear mixed models (LMMs) with a random intercept for neighbourhood were fitted to assess correlation. Residual diagnostics indicated non-normality for all three outcomes. For drain water, the ICC was low (0.06), and the LMM had a higher AIC than the linear model. For produce water, no random-intercept variance could be estimated, indicating no neighbourhood correlation. In contrast, for drinking water, the ICC was higher (0.24) and the random-intercept model reduced the AIC by 47 units
                relative to the linear model. For comparison to the reproduced results, LMMs were bootstrapped for drain and drinking water, whereas the linear model was bootstrapped for produce water, given the absence of a cluster effect. Results for drain and produce water were inferentially reproducible, with a bootstrap standardized difference <0.1; for drinking water, the CI range differed by 0.16 and was not inferentially reproducible.
                Paper 25 This paper examined growth differences between undernourished children who were tuberculin skin test-positive and tuberculin skin test-negative in a disadvantaged Bangladeshi cohort. Statistical methods were clearly described and appropriately
                presented in tables identifying the tests used. Although the paper used some acronyms excessively, it was generally easy to read. The assumptions were not discussed, but non-parametric tests were used for univariate modelling. The authors used pre–post change scores in the linear regression but did not explicitly address independence, and they interpreted the coefficients only in terms of statistical significance. Three primary
                outcomes were analysed using multivariable linear regression: change in length-for-age Z score (LAZ), weight-for-age Z score (WAZ), and weight-for-length Z score (WLZ). The primary independent variable was Tuberculin skin test status, with adjustment for sex and post-intervention age. The models were mainly computationally reproducible, with only minor rounding differences. P-values reported in the paper showed minor rounding discrepancies, consistent with truncation to two decimal places rather than rounding, but these differences did not affect their interpretation. Regression coefficients retained their direction. Some data manipulation was required, but overall, the study was straightforward to reproduce.
                Two of the three models were inferentially reproducible, indicating that the paper was overall not inferentially reproducible. For the LAZ model, the residuals-versus-fitted plot showed a funnel-shaped pattern, suggesting heteroscedasticity, and a wild bootstrap was used as a comparison hwich showed no meaningful differences in standardized coefficients. The Q–Q plot and Kolmogorov–Smirnov test indicated that the residuals were approximately normal, although the Shapiro–Wilk test suggested minor deviations from normality. The covariance ratio flagged potential outliers that could inflate coefficient variances and widen confidence intervals. The study sampling
                method was unclear, and independence could not be tested as there were no identifiers in the dataset. The WAZ and WLZ models exhibited similar assumption issues but did not show evidence of heteroscedasticity, and BCa bootstrapping was used as a comparison for these models. The direction and statistical significance of effects remained consistent between the reproduced and bootstrapped models for all outcomes.
                However, the WLZ model was not inferentially reproducible, with standardized coefficient range differing by 0.11 (14%). In contrast, the LAZ and WAZ models’ differences in standardized coefficients and CI ranges were <0.1, indicating inferential
                reproducibility.
                Paper 27 This study examined metabolites, body weight, and puberty onset in male African lions using repeated-measures ANOVA and linear regression. Reporting was often unclear, particularly regarding which analyses corresponded to which results, making it difficult to determine the total number of regression models fitted. The data were provided in an unformatted Excel file. From the description provided, four linear regression models appear to have been used to examine weight gain over 30 months: one for wild male lions and three for captive males, comprising one overall model and two stratified by follow-up duration (up to 24 months and beyond 24 months). The paper was not computationally reproducible, as only one of these models could be reproduced. The regression model for wild male weight was successfully replicated, although interpretation required recoding the time variable from 0–30 months to 1–31 months to match the reported intercept. The two models examining captive male weight could not be reproduced; the timing was unclear, and the model beyond 24 months was based on only three observations due to missing data. The model for wild lion weights was not inferentially reproduced due to concerns about linearity. A quadratic term for month was therefore considered. Relative to the linear specification, the quadratic model provided a modestly improved fit (AIC lower by 6.2; BIC lower by 4.9) and improved diagnostic behaviour, with clearer residual patterns and attenuation of departures from linearity and normality. The quadratic term captured mild concave-down curvature in the month–weight relationship rather than a strictly linear increase. On balance, the quadratic model was considered better specified and conceptually reasonable and we decided this was a better fitting model.
                Paper 30 This paper is an exploratory analysis of factors associated with bone turnover (bone resorption and formation). A stepwise linear regression approach was reported to explore turnover biomarkers (described by the authors as a “stepping method”);
                however, the direction (forward, backward, or bidirectional) was not explicitly stated, although entry (0.05) and removal (0.10) criteria were specified. The authors checked for normality of continuous variables. Linearity was not discussed, but scatterplots were shown. Regression coefficients were not interpreted, conclusions were made in terms of direction and statistical significance, and there was no discussion of false positives with multiple outcomes. Collinearity was not mentioned. Three outcomes were randomly selected for further evaluation: log-transformed cross-linked telopeptide of type 1
                collagen (lg-CTX-1), tumour necrosis factor (TNF), and log-transformed osteoprotegerin (lg-OPG). This paper was mostly computationally reproduced. The authors gave unadjusted R2 but incorrectly reported them as adjusted R2. There were minor discrepancies in the second and third models attributable to typographic errors in the sex variable related to incorrect decimal point placement. The paper was straightforward to reproduce, as the data were well formatted in SPSS and required no additional data cleaning or reformatting.
                The models assessed for inferential reproducibility were reproducible overall. Only minor violations of model assumptions were identified, including homoscedasticity, linearity, and normality. Statistically significant quadratic terms were observed for the lg-CTX-1 and lg-OPG models; however, diagnostic plots and model fit statistics did not provide substantive support for meaningful non-linear relationships. The study design appeared to use independent obsevations; however, the lg-CTX-1 model showed a significant Durbin–Watson test. Further inspection suggested that this was likely due to post-randomisation sorting of control and intervention cases rather than true autocorrelation. All three models yielded statistically significant Shapiro–Wilk tests for normality; however, Q–Q plots and the Kolmogorov–Smirnov test indicated approximately normally distributed residuals. Mild heteroscedasticity was observed for the TNF and lg-OPG models. Bootstrapped comparisons showed that although some statistics differed by 10% or more, these differences were not meaningful, with
                standardized regression coefficients remaining below 0.1. The direction of effects and statistical significance were consistent between the reproduced and bootstrapped models.
                Paper 32 This paper examined factors associated with country-level grant allocation in the European Union using linear regression. Multivariable modelling appeared consistent with a backward selection approach, although this was not explicitly described. Regression assumptions were assessed using residual diagnostics, and collinearity and outliers were examined; however, influential observations were retained without accompanying sensitivity analyses. The study was not fully computationally reproducible. While the univariate analysis for citations and the final multivariable model were mostly reproducible (with only minor rounding differences), one of the three models could not be reproduced. In particular, the initial multivariable model differed from the reproduced results, despite the final model differing only by the inclusion of DALYs. Investigation of the DALY results identified discrepancies suggesting potential differences in the data. This paper was not inferentially reproducible. The univariate model for citations was successfully reproduced, with standardized coefficients and confidence interval bounds below 0.10, and with consistent coefficient directions and statistical significance. However, the final multivariable model contained a highly influential observation (Cook’s distance = 3.9). The sensitivity analyses indicated substantial changes in effect sizes (The GDP coefficient more than doubled when the outlier was removed) and linearity conclusions. To mitigate the influence of an extreme observation, funding and GDP per capita were log-transformed, while citations were retained on the original scale. As the model was fitted on the log scale, fit statistics (e.g. R2, AIC, BIC) are not directly comparable with those from models on the original scale. Although the influential observation continued to have some impact, its influence was substantially reduced, with Cook’s distance below 0.5. On the log scale, residual diagnostics did not indicate evidence of non-linearity in the association between GDP per capita and funding. Mild heteroscedasticity was observed, robust standard errors were used as a sensitivity analysis.
                Paper 33 This study assessed median nerve cross-sectional area using ultrasonography in patients with essential tremor, compared findings with controls, and examined associations with nerve conduction parameters and tremor severity. ANCOVA and linear regression were identified as the primary statistical methods; however, the authors did not report checks of model assumptions, outliers, or collinearity. Both hands from individual
                patients were included in the analyses without adjustment for repeated measures. Regression coefficients were not meaningfully interpreted beyond their direction. Forward stepwise selection was used for the linear regression models. The results based on the provided data were computationally reproducible. In addition, only a subset of the data was available, with no data provided for the control group. This resulted in one of the three randomised models not being assessed for reproducibility.
                Two models were assessed for inferential reproducibility with carpal inlet as the outcome. The first model included sensory amplitude and age as predictors, and the second included sensory nerve conduction velocity and age. Neither model was inferentially reproducible. The primary limitation was the lack of adjustment for repeated measures, as both left and right hands from each participant were included in the analyses with ICCs>0.3. Evidence of heteroscedasticity was also observed. Consequently, linear mixed models with a wild bootstrap were fitted to account for within-participant clustering and heteroscedasticity. Comparisons of coefficients
                and 95% confidence intervals indicated medium (>0.10) to large (>0.20) differences in the standardized confidence-interval ranges. The direction of the coefficients and the statistical significance were consistent.
                Paper 38 This paper examined associations between vitamin D and seven sleep phenotypes using linear regression, with covariate adjustment via a stepwise modelling approach; however, the direction of selection (forward, backward, or bidirectional) was not reported. Model assumptions were assessed using residual diagnostics, and interpretation focused primarily on effect direction and p-values. Seven multivariable models were identified as the primary analyses. The study was not computationally reproducible. Of the three reported multivariable models, only the Subjective Sleep Quality model could be reproduced. Although the authors provided a detailed description of their statistical methods, reproducibility was limited by the absence
                of derived variables in the provided dataset. More than 30 transformed, centred, and dummy-coded variables had to be recreated, including those subjected to Box–Cox transformations. The remaining two models could not be reproduced due to inconsistencies in reported sample sizes and degrees
                of freedom. The authors reported missing data for only one independent variable, but these were not consistent with that variable. The first model reproduced indicated that stepwise modelling was implemented with full cases rather than listwise deletion. In addition, the Midsleep Time model showed a mismatch between the number of estimated coefficients and the
                degrees of freedom reported in the ANOVA table. Collectively, inconsistencies in derived variables, sample sizes, and model specification rendered two of the three analyses computationally non-reproducible.
                Although the authors provide a detailed description of their modelling strategy, it focuses primarily on statistical procedures and assumptions rather than on the clinical or substantive context of the analysis. While the bootstrapped regression coefficients and p-values were reproduced, the model was not inferentially reproducible. Only the final model was assessed;
                however, the modelling process described in the manuscript was flawed and reflected two key misconceptions about statistical modelling. First, the authors described stepwise variable selection as a solution to multicollinearity. Highly collinear predictors were entered into a stepwise selection procedure, a practice known to produce unstable coefficient estimates and p-values. Because correlated predictors compete to explain
                the same variance, minor changes to the data or model specification can cause variables to enter or exit the model, change coefficient signs, or alter statistical significance. As a result, findings may appear computationally reproduced while remaining inferentially unstable. Collinearity should be addressed before any stepwise selection. Second, categorical predictors
                were dummy-coded and treated as separate independent variables within the selection process, including interaction terms. This resulted in interactions being evaluated independently of their corresponding main effects and multi-level categorical variables being fragmented across multiple tests. Together, these modelling decisions undermine the stability, interpretability, and inferential validity of the reported results.
                Paper 41 This paper examined the association between percentage weight loss and health-related quality of life following a health intervention using multivariable linear regression. Models adjusted for age, comorbidities, change in physical activity, intervention group, and baseline outcome scores. Only the coefficient direction was interpreted, and results were reported solely for the primary predictor of weight change. The normality of
                the raw data was mentioned, but other assumptions, outliers, and collinearity were not discussed. The study was computationally reproducible. All results for the depression model were reproduced, and the sleep-disturbance and fatigue models were mostly reproducible. One p-value discrepancy in the sleep-disturbance model (reported 0.67, reproduced 0.57) was considered typographical and did not affect significance. A minor rounding difference was observed in one fatigue coefficient and R2 was incorrectly reported instead of adjusted R2
                All three models were inferentially reproducible. Residual diagnostics did not indicate major violations of linear model assumptions. A small number of larger standardized residuals were also present, which may contribute to wider confidence intervals but did not materially affect point estimates or inference. Although some statistics differed by 10% or more across models, these differences were not considered meaningful because the corresponding changes in standardized regression coefficients were less than 0.10. The direction of effects and statistical significance remained consistent between the reproduced models and the bootstrapped sensitivity analyses.
                Effect size comparisons between the reproduced and sensitivity results were restricted to directly comparable models, excluding those with specification differences (e.g., non-linearity). Standardized regression coefficients were evaluated using percentage differences in estimates and changes in the confidence interval range, with interpretation primarily focused on the confidence interval range as an indicator of changes in precision rather than on changes in the mean effect alone. Across 97 coefficient ranges evaluated, 30 (31%) exceeded the 10% change threshold. When applying absolute thresholds based on standardized differences, 12 (12%) exceeded a change of 0.1 and 2 (2%) exceeded a change of 0.2. At the paper level, a paper was classified as exceeding a threshold if any coefficient range exceeded it. Nine of 11 (80%) papers exceeded the 10% threshold, while 5 (50%) exceeded the 0.1 threshold and 2 (18%) exceeded the 0.2 threshold, indicating some variation in classification depending on the threshold applied. Although coefficient thresholds were only one component of the reproducibility decision, a standardized difference of 0.1 was used as the threshold for determining reproducibility in this study.

                In contrast to changes in confidence interval width, shifts in statistical significance were uncommon. Of the 97 p-values assessed, we observed only one (1%) shift in significance classification (from significant to non-significant or vice versa), with 40 (41%) remaining statistically significant and 56 (58%) remaining non-significant.

                Despite this relative stability in significance classification, a substantial proportion of models exceeded thresholds for changes in confidence interval width or demonstrated model assumption violations (e.g., linearity or independence) and were classified as not inferentially reproducible.

                Discussion

                Despite the rapid growth in health data and scientific output, inferential reproducibility, the extent to which independent analysts draw similar conclusions from the same data, remains rarely assessed [3]. Our study highlights the practical challenges of evaluating this form of reproducibility, including incomplete reporting of statistical methods, vague model specifications, and limited information on key assumptions [19]. For example, some studies used repeated measures designs, yet the author-supplied datasets lacked participant identifiers, making it impossible to assess for the effect of clustering. This reflects a broader challenge of inferential reproducibility, which requires evaluating a series of complex modelling decisions that should be guided by the study design but are often poorly documented, making the process labour intensive.

                Our results indicate that inferential reproducibility is often compromised when model assumptions are violated. In many cases, confidence intervals around estimates changed substantially, raising concerns about the reliability of the reported effects.

                Given that such findings may inform decisions about which health treatments are incorporated into standard practice, we recommend that studies making claims with potential clinical or policy implications be subject to inferential reproducibility assessments. Others have argued that reproducibility efforts should be strategically focused on high-impact studies, such as those with over 250 citations, to ensure that widely disseminated findings are robust and to direct limited resources toward reproducibility [44]. However, this may be too late as by the time papers are cited this much, they are entrenched in the field; it may be better to focus on highly cited papers in the first year, e.g. 10 to 20. While awareness of reproducibility is growing, meaningful progress will require not only stronger reporting standards, targeted investment, and clearer incentives from journals and funders [45, 46], but also improved training in statistical reasoning and greater emphasis on evaluating model assumptions. These elements are essential to support researchers in making informed modelling decisions and producing reproducible, valid inferences.

                Although automation is often proposed as a way to scale reproducibility checks, its feasibility depends on the aspect of reproducibility being assessed [47]. For inferential reproducibility, many modelling decisions require context-specific judgment that cannot be reliably captured through generic rules. While automation may help flag major assumption violations, such as extreme non-normality or heteroscedasticity, other decisions, such as how to address non-normality or whether to include non-linear terms, should not be based solely on the data, but guided by the study design and subject-matter considerations. These judgments require a clear understanding of both the model and the data. For this reason, any automated approach must be supplemented by expert human review, which can interpret assumptions in context and determine whether the chosen methods are appropriate.

                Comparison to other studies

                A well-known case highlighting the consequences of inferential failure comes from economics. Reinhart and Rogoff [48] reported that countries with high debt-to-GDP ratios experienced lower economic growth, a finding that influenced global austerity policies. A spreadsheet error later uncovered by a graduate student revealed major flaws in the analysis [49]. These austerity measures led to reduced public spending in several countries, including cuts to health and social care budgets. In the UK, for example, such cuts were linked to excess deaths and widening health inequalities [50]. This case illustrates how analytical errors can shape policy with real-world health consequences, reinforcing the importance of reproducibility in evidence-based decision-making.

                It is important to note that our study focused on statistical assumptions and did not consider variable selection, which can also affect reproducibility. This is clearly demonstrated by Silberzahn et al. [51], who asked 29 research teams to analyse the same dataset and research question, these types of studies are known as “Many Analyst” projects. Despite using the same data, teams applied 21 different covariate sets and a range of analytic methods, including linear regression, multilevel models, and Bayesian approaches, producing a large range in effect sizes. While 69% of teams reported statistically significant results, their conclusions varied, and choices regarding model type and handling of non-independence (e.g., variance components, clustered standard errors, or fixed effects) influenced the outcomes. The authors found that statistical expertise did not fully account for this variability, concluding that even well-intentioned analysts may reach different conclusions due to methodological flexibility in complex datasets.

                Another Many Analyst project conducted by Gould et al. [52] showed how analytical decisions can substantially influence research findings, even when using the same data.

                In their study, 174 analyst teams analysed two ecological datasets on blue tit nestling growth and Eucalyptus seedling recruitment. The resulting effect sizes varied widely, with some teams finding strong effects, others near zero, and some even in the opposite direction. These findings further highlight how different analytical choices can lead to markedly different conclusions.

                These many-analyst studies, like our findings, highlight that reproducibility is not just about data access but also how analytical decisions, such as model structure and handling of assumptions, can substantially influence conclusions [53]. While our study focused on inferential reproducibility, others have argued that reproducibility alone is insufficient. Munafo and Smith [54] advocate for triangulation using multiple, independent approaches to answer the same question, since replication can still reinforce flawed assumptions [55]. Although debate continues about the scale of the reproducibility “crisis,” there is broad agreement that improving methodological rigour and statistical reasoning is essential [56]. Fanelli [57] proposes reframing the issue as an opportunity for researchers to take greater responsibility for research quality.

                Interpreting Inferential Reproducibility: Beyond Statistical Significance

                Our study contributes to the limited literature on inferential reproducibility by examining whether violations of linear regression assumptions influence study conclusions. We identified a number of papers in which linear regression was applied inappropriately, resulting in confidence intervals that were frequently too narrow. This underestimation of uncertainty suggests reduced precision and overstated confidence in the reported effects. In several cases, appropriately addressing these violations required fitting alternative models, such as replacing a linear model with a Tweedie or Inverse Gaussian model. A likely contributor to poor inferential reproducibility is limited understanding of model assumptions. Many researchers either do not formally assess assumptions or rely on statistical tests without recognising their limitations [9]. In our previous work, drawing on the larger random sample of 95 papers, we found that the most common misconception was assessing normality of the outcome rather than the residuals [14]. Of the 28 author teams who reported checking normality, only five evaluated residuals, highlighting a substantial gap between theoretical knowledge and applied practice and reflecting persistent misunderstandings about the role and interpretation of regression assumptions.

                One example from the current study where a violation of normality should have prompted reconsideration of the outcome distribution is found in paper 22. While two of the three models showed only mild deviations from normality of residuals, one showed a substantial deviation. Inspection of the outcome data revealed that nearly a quarter of the values were zero. Several modelling approaches are available for handling zero-inflated data. Two types of models address structural zeros, while a third treats zeros as part of a continuous mixture [22]. For instance, imagine a study measuring post-treatment quality-of-life scores on a 0–100 scale, where zero represents no impairment. Thirty percent of patients report a score of zero, while the remaining 70% have scores between 1 and 100. This pattern fits a Tweedie distribution, a flexible family that includes the Poisson and Gamma distributions as special cases and accommodates a mean–variance relationship in which the variance increases with the mean. Alternatively, a hurdle model is appropriate when zeros arise solely from a structural mechanism, for example, when a patient did not experience the condition being assessed and therefore could not have an impairment. A zero-inflated scenario occurs when both structural zeros (patients who could not have impairment) and sampling zeros (patients who could have impairment but reported none) are present in the same dataset. The complexity of model choice and interpretation highlights the importance of consulting an experienced statistician, as papers involving statisticians tend to have fewer errors, including inappropriate test selection [58].

                Assessing whether model assumptions are valid is often more similar to a mixed-methods approach than a purely quantitative process. As in qualitative research, where triangulation of multiple data sources can assess whether they tell a consistent story, assumption checking should integrate both quantitative indicators and contextual understanding. For example, knowledge of the study design is essential when assessing independence. Quantitatively, the model can be examined for the level of correlation between measurements using the intra-class correlation coefficient (ICC); the AIC can inform model complexity; and effect sizes and confidence intervals can be compared across models to assess changes. While one of these may be the dominant consideration in some cases, a full evaluation typically requires combining all sources of information to make an informed decision. For instance, when the ICC is relatively low, researchers might still choose to use a linear mixed model rather than a simple linear regression to remain consistent with the study design. However, a high ICC indicates that simple linear regression is inappropriate. We observed this in paper 16, where both models assessed had high ICCs (0.26, 0.35), indicating substantial within-cluster correlation.

                In our study, p-value classifications were relatively stable, with few changes from statistically significant to non-significant and vice versa. However, this finding should be interpreted in context and cannot be directly compared to studies reporting inconsistencies in p-values [59], as our analyses were restricted to models that had already been found to be computationally reproducible. Some may argue that if statistical significance remains unchanged, there is little cause for concern. Yet statistical significance is a poor arbiter of importance. Small studies often remain non-significant despite meaningful shifts in effect magnitude or precision, whereas large studies can retain statistical significance even when changes are trivial or clinically irrelevant. Our assessment showed that many authors framed their conclusions primarily in terms of statistical significance rather than effect magnitude or precision [15]. Although significance categories were relatively stable, confidence interval widths were frequently not reproducible. This indicates that unchanged statistical significance does not imply inferential stability. Precision is particularly important when estimating clinically relevant change, as the width of the confidence interval helps define the range of plausible effects and therefore the degree of certainty when weighing potential benefits against risks for new treatments.

                To evaluate changes in effect size, we examined three thresholds. A threshold of 10% change in estimates flagged many inconsequential differences, and we do not recommend this approach for assessing inferential reproducibility. In contrast, absolute thresholds of 0.1 and 0.2 for changes in standardized regression coefficients performed more appropriately. The lower threshold (0.1) was more consistent with cases in which model violations were present. For example, in one instance with an ICC of 0.26, the change in the range of the standardized confidence interval was approximately 0.15. A simulation study to evaluate the utility of these thresholds under controlled conditions would be a valuable next step.

                Challenges in Assumption Checking

                An important consideration in evaluating model assumptions is that the process requires judgment and visualisation, rather than relying solely on strict, hypothesis-testing-based rules, which can sometimes be misleading [60]. Even statisticians may reasonably disagree on the best approach. Multi-analyst studies have demonstrated how different analytic decisions can yield varying results [52, 51]. It is also important to note that residuals are estimates of the true errors [61]. Overly flexible approaches, such as fitting Generalized Additive Models (GAMs) to residuals can overfit noise, mistaking random variation for structure [62, 63]. While minor violations of assumptions may not always justify alternative models, we observed substantial violations in several papers in which linear regression was inappropriate yet used without question.

                These challenges are compounded by the way statistics is commonly taught. Introductory courses often present methods such as t-tests, ANOVA, and linear regression in a simplified, rule-based manner, typically applied to clean, “well-behaved” data [64]. These models are rarely introduced as part of a unified framework, namely the General Linear Model (GLM), leading students to view them as unrelated procedures rather than variations of the same underlying structure [14]. Measures such as R2 are frequently taught as key indicators of model quality, despite offering limited insight into model appropriateness or reliability. Broader tools for model assessment, such as the Akaike Information Criterion (AIC), may be less often introduced, leaving researchers with a limited understanding of how to evaluate fit or compare models.

                Moreover, violations of statistical assumptions are often treated as technical checklist items rather than potential indicators of model misspecification. For example, apparent non-normality of residuals may reflect omitted variables or an incorrectly specified relationship between variables rather than a problem with the outcome distribution.

                Similarly, outlier handling in some textbooks and basic statistics courses is often taught using simple cut-offs (e.g., three standard deviations from the mean), rather than through influence diagnostics such as Cook’s distance, DFFITS, or DFBETAS. While these measures focus on changes in regression coefficients, diagnostics that assess the precision of estimates, such as COVRATIO, which evaluates the inflation of standard errors and the potential widening of confidence intervals, are often neglected, leaving applied researchers without a complete picture of how outliers affect inference [25]. Many textbooks also provide conflicting advice on whether to exclude outliers, often without evaluating the appropriateness of removal based on influence or its impact on model validity [65].

                Although excluding outliers without strong justification is generally inappropriate, influential observations should be explicitly examined in sensitivity analyses to assess their impact on estimates and inference. Approaches such as robust regression and bootstrapping are commonly used in this context. Robust regression downweights influential observations during model fitting, thereby reducing their impact on coefficient estimates and associated inference. Bootstrapping can reduce reliance on strict distributional assumptions by resampling the observed data; however, it assumes that the fitted model is correctly specified and that influential observations are representative of the underlying data-generating process. Neither approach addresses fundamental model misspecification, such as incorrectly modelling the relationship between variables.

                When such influential observations were identified, their impact was evaluated to determine whether they distorted parameter estimates or reflected broader model assumption violations. In paper 46, we resolved linearity and outlier issues by log-transforming both the predictor and the outcome, but high-leverage observations continued to have a disproportionate influence on the fitted slope. Robust regression was then applied to the log–log model to down-weight these observations during estimation, producing coefficients that more closely reflected the central trend of the data. A similar issue was observed in paper 32, log-transformation similarly was used to resolved departures from linearity and substantially reduced the influence of extreme observations. However, residual heteroscedasticity persisted, indicating a violation of the constant-variance assumption. Robust standard errors were therefore applied to the log–log model to obtain valid inference without altering the point estimates. These examples highlight the distinct roles of robust-based methods: robust regression is appropriate when influential observations distort parameter estimates, whereas robust standard errors are appropriate when variance assumptions are violated but the model is otherwise adequate.

                Some argue that many authors do check model assumptions but, due to space constraints in journals, do not report the details or even mention that checks were conducted [66]. To our knowledge, no empirical research has formally tested this claim. Our findings suggest that the proportion of papers clearly addressing assumptions is low. While some evidence suggests that a few authors assessed assumption violations, these violations were rarely reported in detail. For example, one paper stated that assumptions were checked and found to have only mild violations that did not substantially affect results. However, many of the papers we assessed for inferential reproducibility reported only limited, and often incorrect, assumption checks, despite major violations that substantially altered the precision of the reported effects.

                Issues of assumption checking also intersect with broader concerns about modelling choices and reproducibility, where software defaults may shape analytic decisions. One example we observed is from paper 22, in which ethnicity was incorrectly modelled as a continuous variable. In SPSS, unless categorical variables are explicitly dummy coded, they are treated as continuous in linear regression models. Similar risks arise when stepwise regression is applied to dummy-coded categorical variables and interaction terms, where software may treat individual terms as independent. A possible example is from paper 38, where substantial effort was devoted to testing assumptions of their interaction model; yet backward stepwise regression was subsequently used, resulting in the elimination of main effects involved in the interaction and components of categorical variables. SPSS also includes a function called Univariate General Linear Model, which performs linear regression within the broader General Linear Model framework. This function automatically dummy-codes categorical variables, which researchers may not realise given its name and setup, although it does not perform variable selection. These examples highlight a key point: understanding your software’s default settings is essential. Seemingly minor choices, such as drop-down selections or default parameterisations, can substantially change the interpretation of results. Without adequate statistical training, researchers risk producing results that are computationally reproducible but inferentially misleading.

                Our findings highlight the need for better statistical education that reflects the complexity and judgment involved in modelling. Qualitative evidence suggests that researchers often interpret regression coefficients through context dependent reasoning rather than as purely mechanical statistical outputs [67], reinforcing the importance of conceptual understanding over rule based application. Some authors in our sample made genuine efforts to satisfy regression assumptions, but this sometimes led to models that were difficult to interpret and overemphasised statistical significance. For example, transformations were sometimes applied and not back-transformed, limiting interpretability on the original scale. In other cases, transformations appeared to be motivated primarily by a desire to “pass” normality tests, even when deviations were unlikely to be practically meaningful. Teaching statistics within a broader modelling framework, one that emphasises critical thinking over rigid rules, would better prepare researchers to apply methods appropriately, consider alternatives such as bootstrapping or other distributions, and prioritise interpretability and practical relevance over arbitrary thresholds.

                General recommendations for improving inferential reproducibility

                • Consider the underlying data-generating mechanism (e.g. continuous, count, or proportions) when selecting analytical approaches. Provide exploratory visualisations in the Supplementary Information, such as scatterplots, boxplots, or other appropriate plots, to illustrate key relationships between variables included in the analyses.
                • Clearly describe the study design, explicitly stating whether the study is observational or experimental, whether randomisation was used, and whether there is temporal ordering, such as in a baseline and follow-up design.
                • Provide a unique observation identifier (ID) in the dataset to ensure each individual is uniquely and consistently identified.
                • Consider whether the study design includes clustering (e.g. by site, school, hospital, or family). If clustering is present, include the relevant clustering variables in the dataset.
                • Clearly describe and justify the modelling strategy, including the rationale for inclusion of confounders, mediators, and interaction terms, and whether a formal model-building framework was used.
                • If a formal modelling process was used, the variable selection and/or removal procedures should be described in detail, including all decision thresholds and criteria. Authors should also report the number and type of candidate variables initially considered for inclusion, rather than only listing those retained in the final model.
                • Evaluate the assumptions of linear regression models by examining residual diagnostics, identifying influential observations, and assessing multicollinearity among predictors. Where assumptions are violated, apply and report appropriate remedial methods.
                • If model assumptions are violated, consider fitting alternative models that better reflect the underlying data structure (e.g. non-linear specifications or models with different distributional assumptions). Use appropriate model-fit statistics, such as R2, Akaike Information Criterion (AIC), or Bayesian Information Criterion (BIC), to compare competing models and justify the final model choice.
                • Assess heteroscedasticity explicitly and, where present, consider using robust standard errors or alternative modelling strategies to ensure valid statistical inference. When systematic variance differences are evident (e.g., differing variances across groups or increasing variability with a predictor), weighted least squares may be considered with appropriate justification.
                • Identify influential observations using appropriate influence diagnostics (e.g. Cook’s distance, DFBETAs, DFFITS, and covariance ratios), and evaluate their impact on coefficient estimates and statistical inference. Where influential observations materially affect results, conduct sensitivity analyses, such as applying robust regression methods to reduce their influence.
                • Transformations may be used to address violations of linearity or homoscedasticity, but should be applied only where they meaningfully improve model adequacy while retaining interpretability. For log-transformed models, back-transformation is often required to present results on the original scale, such as geometric means or ratios (interpreted as percentage or fold changes). Researchers should consider whether a transformation is necessary and whether the resulting parameter estimates correspond to the effect of interest, recognising that transformations (e.g. log scale) change interpretation from additive differences to multiplicative or proportional effects.
                • Bootstrapping can provide standard errors and confidence intervals without relying on the assumption that residuals are normally distributed. However, it does not address model misspecification, such as non-linearity. When bootstrapping is used, the resampling scheme (e.g., residual, case, wild, or cluster bootstrap), the number of resamples, and the model coefficients and confidence intervals should be clearly reported.
                • Report the details of assumption checking in the Supplementary Information, including diagnostic plots and any sensitivity analyses undertaken.
                • Provide a clear description of the extent and patterns of missing data, including the assumed missing data mechanism (e.g., missing completely at random, missing at random, or missing not at random). Authors should detail the analytical approach used to address missingness (e.g., complete case analysis, single or multiple imputation, inverse probability weighting), along with any sensitivity analyses conducted to assess the robustness of results.
                • Share code, underlying data, and a comprehensive data dictionary with a well-documented workflow to clarify modelling decisions and support inferential reproducibility.

                  Limitations

                  While every effort was made to faithfully reproduce published analyses using the methods described in the original articles, evaluating statistical assumptions is inherently subjective. Contextual knowledge, often unavailable to external analysts, can influence judgments about the severity and relevance of assumption violations.

                  Consequently, reasonable differences in interpretation and analytic decisions may arise among statisticians and researchers. In addition, implicit methodological details, such as data-handling procedures or variable coding, were often not fully reported, which may affect reproducibility. Nonetheless, we maintain that research articles should provide sufficient detail for any qualified statistician to independently reproduce the analyses.

                  Our results should be interpreted with caution, as we only assessed inferential reproducibility in papers that were computationally reproducible, a subset that may disproportionately represent methodologically skilled authors. The primary author evaluated inferential reproducibility, with input from co-authors as required. Although our protocol initially specified a 10% change in coefficients as a threshold, this was often not meaningful when coefficients were close to zero. To address this, additional comparisons were made using cut-offs of 0.1 and 0.2 for standardized regression coefficients.

                  As this was an exploratory study, the methodology was refined as challenges were identified; the original study protocol is available on GitHub [29]. Overall, our findings align with previous research documenting concerns about inferential reproducibility in scientific research [3] and with limited or inconsistent reporting of model assumptions [68].

                  Conclusions

                  Assessing inferential reproducibility is inherently complex, as it requires judgment about whether modelling choices and assumptions are appropriate for the study design and research context. While major issues, such as extreme non-normality or severe heteroscedasticity, are relatively straightforward to identify, more nuanced decisions, such as whether to include non-linear terms or account for zero inflation, require contextual understanding and cannot be reduced to simple rules without risking misinterpretation.

                  The limited reporting of assumption checks observed in our study suggests that modelling decisions may not always be explicitly justified. In some instances, model specifications appeared consistent with default software settings rather than decisions clearly aligned with the research question. Linear regression models were frequently used without explicit assessment of key assumptions or consideration of whether it was the most appropriate modelling approach.

                  To help guide researchers, we have developed general recommendations for inferential reproducibility informed by issues identified across the studies examined. The recommendations draw directly on recurring problems observed in the reporting and interpretation of statistical analyses and are intended to provide practical guidance to improve the transparency and consistency of inference. We recommend using these recommendations when developing statistical analysis protocols, as they promote explicit, up-front consideration of modelling assumptions, analytical decisions, and planned sensitivity analyses, and revisiting them during the analysis stage to ensure these considerations are implemented and transparently reported.

                  In our experience, introductory statistics courses are often taught in ways that encourage simplified responses to violations of assumptions, such as defaulting to nonparametric tests. This can leave researchers insufficiently prepared to manage the complexities of statistical modelling. We recommend introducing the foundational principle that “everything is a regression” [69], whereby t-tests, ANOVA, and linear regression are understood as special cases of the general linear model. Extending this framework to include link functions and Generalised Linear Models would help students develop a deeper understanding of modelling from the outset. Although it is not feasible to teach all statistical techniques in a single introductory course, conveying this broader conceptual framework is particularly important given that many health professionals complete only one formal course in statistics during their university degrees.

                  As analyses become more complex, researchers should be encouraged and supported to consult with qualified statisticians, ensuring that modelling choices are appropriate, assumptions are adequately evaluated, and results are valid and interpretable. However, in many hospital and health service contexts, access to this expertise is limited. This is particularly concerning given the potential consequences of low-quality studies in health settings, where inappropriate modelling or unchecked assumptions can lead to misleading conclusions. A stronger understanding of how decisions such as model parameterisation and distributional assumptions influence conclusions is essential for informed decision-making. Reproducibility is more than a technical exercise; analyses may computationally reproduce yet still produce misleading conclusions if they rely on flawed modelling choices.

                  While minor violations of statistical assumptions should be discussed in manuscripts and comment sections, situations involving gross violations, such as changes to substantive conclusions or the use of inappropriate models, warrant a formal correction. At present, however, practices such as sharing data or documenting model assumptions are often treated as optional or even risky, with the potential to result in negative professional consequences if errors are discovered [70]. A further barrier is the perception that acknowledging mistakes or issuing corrections damages a researcher’s and journals’s reputation. Journals and publishers could help reduce these barriers by embedding reproducibility into standard expectations and by making the correction process straightforward and routine. Institutions also have a responsibility to foster a culture in which correcting errors is expected and valued, positioning transparency and correction as markers of professional integrity rather than sources of stigma.

                  Journals also have a crucial role in strengthening peer review by engaging qualified statistical reviewers and enforcing compliance with data availability requirements. Some disciplines have adopted more structured editorial models to support reproducibility; for example, the American Economic Association employs a dedicated data editor to review replication materials prior to publication [71]. Similar mechanisms in health research could improve scrutiny of modelling decisions and enhance inferential reproducibility.

                  Our findings reinforce ongoing calls for stronger methodological and statistical rigour in health research. While computational reproducibility is valuable, it is insufficient to ensure scientific reliability. Inferential reproducibility is essential for understanding the magnitude and precision of results and depends on appropriate modelling decisions, interpretation, and assessment of assumptions. Meaningful improvements will require improved statistical literacy, practical training in model diagnostics, and systemic changes that support high-quality research practices. Reproducibility must be understood not merely as the ability to rerun code, but as the capacity to justify, evaluate, and defend statistical decisions in a transparent and principled manner.

                  Strengthening inferential reproducibility is therefore central to improving the credibility, interpretability, and practical value of health research.

                  Data Availability

                  The data and a reproducible R Quarto file used to produce this paper, including tables, figures, and code, have been stored in a GitHub repository and can be cited using the Zenodo doi.

                  https://github.com/Lee-V-Jones/Reproducibility

                  https://doi.org/10.5281/zenodo.19448969

                  Acknowledgements

                  We acknowledge Adrian llich for providing bioinformatics support and all the statisticians, including those acknowledged by name and others who chose to remain anonymous, who generously contributed their time to reviewing papers for this study, including: Ingrid Aulike, Peter Baker, Brigid Betz-Stablein, Enrique Bustamante, Taya Collyer, Susanna Cramb, Alanah Cronin, Laura Delaney, Zoe Dettrick, Eralda Gjika Dhamo, Des FitzGerald, Peter Geelan-Small, Edward Gosden, Alison Griffin, Jenine Harris, Cameron Hurst, Kyle James, Helen Johnson, Jessica Kasza, Karen Lamb, Stacey Llewellyn, James Martin, Miranda Mortlock, Satomi Okano, Alan Rigby, Michael Steele, Megan Steele, Jacqueline Thompson, Simon Turner, Michael Waller, Kevin Wang, Jace Warren, Natasha Weaver, Lachlan Webb, and Janet Williams.

                  References

                  1. Cobey KD, Ebrahimzadeh S, Page MJ, Thibault RT, Nguyen PY, Abu-Dalfa F, et al. Biomedical researchers’ perspectives on the reproducibility of research. PLoS biology. 2024;22(11):e3002870.

                  2. Baker M. Is there a reproducibility crisis? A Nature survey lifts the lid on how researchers view the’crisis rocking science and what they think will help. Nature. 2016;533(7604):452455.

                  3. Goodman SN, Fanelli D, Ioannidis JP. What does research reproducibility mean? Science translational medicine. 2016;8(341):341ps12341ps12.

                  4. Hardwicke TE, Wallach JD, Kidwell MC, Bendixen T, Crüwell S, Ioannidis JP. An empirical assessment of transparency and reproducibility-related research practices in the social sciences (2014–2017). Royal Society open science. 2020;7(2):190806.

                  5. Fidler F, Wilcox J. Reproducibility of Scientific Results; 2018. https://plato.stanford.edu/archives/win2018/entries/scientific-reproducibility/.

                  6. Glasziou P, Altman DG, Bossuyt P, Boutron I, Clarke M, Julious S, et al. Reducing waste from incomplete or unusable reports of biomedical research. The Lancet. 2014;383(9913):267276.

                  7. Gelman A, Loken E. The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time. Department of Statistics, Columbia University. 2013;.

                  8. Poorolajal J, Cheraghi Z, Irani AD, Rezaeian S. Quality of cohort studies reporting post the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement. Epidemiology and health. 2011;33.

                  9. Thiese MS, Arnold ZC, Walker SD. The misuse and abuse of statistics in biomedical research. Biochemia medica: Biochemia medica. 2015;25(1):511.

                  10. Boutron I, Ravaud P. Misrepresentation and distortion of research in biomedical literature. Proceedings of the National Academy of Sciences. 2018;115(11):26132619.

                  11. Arnold LD, Braganza M, Salih R, Colditz GA. Statistical trends in the Journal of the American Medical Association and implications for training across the continuum of medical education. PLoS One. 2013;8(10):e77301.

                  12. Ernst AF, Albers CJ. Regression assumptions in clinical psychology research practice a systematic review of common misconceptions. PeerJ. 2017;5:e3323.

                  14. Jones L, Barnett A, Vagenas D. Common misconceptions held by health researchers when interpreting linear regression assumptions: a cross-sectional study. PLoS One. 2025;20(6):e0299617.

                  15. Jones L, Barnett A, Vagenas D. Linear regression reporting practices for health researchers, a cross-sectional meta-research study. PloS one. 2025;20(3):e0305150.

                  16. Jones L, Barnett A, Hartel G, Vagenas D. Challenges in the Computational Reproducibility of Linear Regression Analyses: An Empirical Study. medRxiv. 2026; p. 121.

                  17. Karimi-Maleh H, Dragoi EN, Lichtfouse E. How the COVID-19 pandemic has changed research? Environmental Chemistry Letters. 2023;21(5):24712474.

                  19. Jones L. Reproducibility; 2026. Available from: https://github.com/Lee-V-Jones/Reproducibility.

                  20. Breheny P, Burchett W. Visualization of Regression Models Using visreg. The R Journal. 2017;9(2):5671.

                  21. Fox J, Weisberg S. An R Companion to Applied Regression. 3rd ed. Thousand Oaks CA: Sage; 2019. Available from: https://www.john-fox.ca/Companion/.

                  22. Zuur AF, Ieno EN, Walker NJ, Saveliev AA, Smith GM, et al. Mixed effects models and extensions in ecology with R. vol. 574. Springer; 2009.

                  23. Zeileis A, Hothorn T. Diagnostic Checking in Regression Relationships. R News. 2002;2(3):710.

                  24. Hebbali A. olsrr: Tools for Building OLS Regression Models; 2024. Available from: https://CRAN.R-project.org/package=olsrr.

                  25. Belsley DA, Kuh E, Welsch RE. Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. New York: John Wiley & Sons; 1980.

                  26. R Core Team. R: A Language and Environment for Statistical Computing; 2024. Available from: https://www.R-project.org/.

                  27. Selker R, Love J, Dropmann D. jmv: The ‘jamovi’ Analyses; 2024. Available from: https://CRAN.R-project.org/package=jmv.

                  28. Zuur AF, Ieno EN, Elphick CS. A protocol for data exploration to avoid common statistical problems. Methods in ecology and evolution. 2010;1(1):314.

                  29. Jones L. Lee-V-Jones/statistical-quality: Protocol; 2024. Available from: doi:10.5281/zenodo.10620146.

                  30. Harris JK, Wondmeneh SB, Zhao Y, Leider JP. Examining the reproducibility of 6 published studies in public health services and systems research. Journal of Public Health Management and Practice. 2019;25(2):128136.

                  31. Laurinavichyute A, Yadav H, Vasishth S. Share the code, not just the data: A case study of the reproducibility of articles published in the Journal of Memory and Language under the open data policy. Journal of Memory and Language. 2022;125:104332.

                  32. Cohen J. Statistical power analysis for the behavioral sciences. routledge; 2013.

                  33. Bates D, Mächler M, Bolker B, Walker S. Fitting Linear Mixed-Effects Models Using lme4. Journal of Statistical Software. 2015;67(1):148. doi:10.18637/jss.v067.i01.

                  34. Lüdecke D, Ben-Shachar MS, Patil I, Waggoner P, Makowski D. performance: An R Package for Assessment, Comparison and Testing of Statistical Models. Journal of Open Source Software. 2021;6(60):3139. doi:10.21105/joss.03139.

                  35. Heyman M. lmboot: Bootstrap in Linear Models; 2019. Available from: https://CRAN.R-project.org/package=lmboot.

                  36. Hartig F. DHARMa: residual diagnostics for hierarchical (multi-level/mixed) regression models; 2016.

                  37. Mammen E. Bootstrap and Wild Bootstrap for High Dimensional Linear Models. The Annals of Statistics. 1993;21(1):255285.

                  38. Hastie T, Tibshirani R, Wainwright M. Statistical learning with sparsity. Monographs on statistics and applied probability. 2015;143(143):8.

                  39. Burnham KP, Anderson DR. Model selection and multimodel inference: a practical information-theoretic approach. Springer; 2002.

                  40. Kass RE, Raftery AE. Bayes factors. Journal of the american statistical association. 1995;90(430):773795.

                  41. Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. Boca Raton, FL: Chapman and Hall/CRC; 2013.

                  42. Goldstein H, Browne W, Rasbash J. Partitioning variation in multilevel models for binary responses. Understanding Statistics. 2002;1(4):223231.

                  43. Vehtari A, Gelman A, Simpson D, Carpenter B, Bürkner PC. Rank-normalization, folding, and localization: An improved [ineq] for assessing convergence of MCMC (with discussion). Bayesian Analysis. 2021;16(2):667718.

                  44. Gelman A, King A. Our proposal for scheduled post-publication review: “Even if each review took twice the effort of the average pre-publication review, our system would add only 1 percent to the total reviewing effort, while providing important perspectives on papers representing more than one-quarter of the citations received by these influential journals.”; 2025. https://statmodeling.stat.columbia.edu/2025/02/26/pp/.

                  45. Hardwicke TE, Mathur MB, MacDonald K, Nilsonne G, Banks GC, Kidwell MC, et al. Data availability, reusability, and analytic reproducibility: Evaluating the impact of a mandatory open data policy at the journal Cognition. Royal Society open science. 2018;5(8):180448.

                  46. Stevens JR. Replicability and reproducibility in comparative psychology. Frontiers in psychology. 2017;8:862.

                  47. Han H. Challenges of reproducible AI in biomedical data science. BMC Medical Genomics. 2025;18(Suppl 1):8.

                  48. Reinhart CM, Rogoff KS. Growth in a Time of Debt. American economic review. 2010;100(2):57378.

                  49. Leek JT, Jager LR. Is most published research really false?Annual Review of Statistics and Its Application. 2017;4(1):109122.

                  50. Martin S, Longo F, Lomas J, Claxton K. Causal impact of social care, public health and healthcare expenditure on mortality in England: cross-sectional evidence for 2013/2014. BMJ open. 2021;11(10):e046417.

                  51. Silberzahn R, Uhlmann EL, Martin DP, Anselmi P, Aust F, Awtrey E, et al. Many analysts, one data set: Making transparent how variations in analytic choices affect results. Advances in Methods and Practices in Psychological Science. 2018;1(3):337356.

                  52. Gould E, Fraser HS, Parker TH, Nakagawa S, Griffith SC, Vesk PA, et al. Same data, different analysts: variation in effect sizes due to analytical decisions in ecology and evolutionary biology. BMC Biology. 2025;23:35. doi:10.1186/s12915-024-02101-x.

                  53. Schweinsberg M, Feldman M, Staub N, van den Akker OR, van Aert RC, Van Assen MA, et al. Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis. Organizational Behavior and Human Decision Processes. 2021;165:228249.

                  54. Munafò MR, Davey Smith G. Robust research needs many lines of evidence. Nature. 2018;553(7689):399401. doi:10.1038/d41586-018-01023-3.

                  55. Patil P, Peng RD, Leek JT. What should researchers expect when they replicate studies? A statistical view of replicability in psychological science. Perspectives on Psychological Science. 2016;11(4):539544.

                  56. Wass MN, Ray L, Michaelis M. Understanding of researcher behavior is required to improve data reliability. GigaScience. 2019;8(5):giz017.

                  57. Fanelli D. Opinion: Is science really facing a reproducibility crisis, and do we need it to? Proceedings of the National Academy of Sciences. 2018;115(11):26282631.

                  58. Woodie BR, Freking JA, Jones GM, Porter J, Fleischer SE, Pauley AG, et al. Frequency of statistical mistakes and associated article characteristics: a cross-sectional analysis of dermatology journals. Dermatology Online Journal. 2025;31(3).

                  59. Nuijten MB, Hartgerink CH, van Assen MA, Epskamp S, Wicherts JM. The prevalence of statistical reporting errors in psychology (1985–2013). Behavior research methods. 2016;48(4):12051226.

                  60. Shatz I. Assumption-checking rather than (just) testing: The importance of visualization and effect size in statistical diagnostics. Behavior Research Methods. 2024;56(2):826845.

                  61. Cook RD, Weisberg S. Residuals and Influence in Regression. New York: Chapman and Hall; 1982.

                  62. Barnett A, Stephen D, Huang C, Wolkewitz M. Time series models of environmental exposures: good predictions or good understanding. Environmental Research. 2017;154:222225.

                  63. Wood SN. Generalized additive models: an introduction with R. chapman and hall/CRC; 2017.

                  64. Moore DS, McCabe GP, Craig BA, Cook LC. The Basic Practice of Statistics. 9th ed. New York, NY: Macmillan Learning; 2017.

                  65. Aguinis H, Gottfredson RK, Joo H. Best-practice recommendations for defining, identifying, and handling outliers. Organizational research methods. 2013;16(2):270301.

                  66. Hu Y, Plonsky L. Statistical assumptions in L2 research: A systematic review. Second Language Research. 2021;37(1):171184.

                  67. Collyer TA. The eye of the beholder: how do public health researchers interpret regression coefficients? A qualitative study. BMC public health. 2024;24(1):10.

                  68. Valentine KD, Buchanan EM, Cunningham A, Hopke T, Wikowsky A, Wilson H. Have psychologists increased reporting of outliers in response to the reproducibility crisis? Social and Personality Psychology Compass. 2021;15(5):e12591.

                  69. Hannay K. Everything is Just a Regression; 2020. Available from: https://medium.com/data-science/everything-is-just-a-regression-5a3bf22c459c.

                  70. Gomes DG, Pottier P, Crystal-Ornelas R, Hudgins EJ, Foroughirad V, Sánchez-Reyes LL, et al. Why don’t we share data and code? Perceived barriers and benefits to public archiving practices. Proceedings of the Royal Society B. 2022;289(1987):20221113.

                  71. Vilhuber L, Son HH, Welch M, Wasser DN, Darisse M. Teaching for large-scale reproducibility verification. Journal of Statistics and Data Science Education. 2022;30(3):274281.

                  Editors

                  Kathryn Zeiler
                  Editor-in-Chief

                  Wolfgang Kaltenbrunner
                  Handling Editor

                  Editorial assessment

                  by Wolfgang Kaltenbrunner

                  DOI: 10.70744/MetaROR.418.1.ea

                  The manuscript presents an empirical assessment of the inferential reproducibility of linear regression analyses in health and biomedical research. Both reviewers agree that it addresses a timely and important question within reproducibility research and highlight the study’s methodological transparency, detailed diagnostic framework, and strong pedagogical value. They note that the paper makes a valuable contribution by moving beyond computational reproducibility to examine how violations of statistical assumptions may affect inferential conclusions and by providing practical guidance for researchers, reviewers, and students.

                  At the same time, both reviews identify areas where the manuscript could be strengthened. A recurring theme concerns the conceptual framing of inferential reproducibility, including its distinction from robustness and sensitivity analysis, the justification of inferential thresholds, and the role of analyst judgement in assessing reproducibility. Reviewers also encourage a more explicit discussion of the exploratory nature of the findings given the relatively small inferential reproducibility sample, the robustness of linear regression to moderate assumption violations, and the subjectivity involved in selecting and evaluating alternative models. Additional suggestions include improving access to supporting materials, clarifying terminology and statistical acronyms, and further contextualizing the study’s findings and limitations. Overall, the reviewers and I agree that this is a valuable and constructive contribution to the meta-research literature that would benefit from clarification and refinement of several conceptual and interpretive aspects.

                  Competing interests: None.

                  Peer review 1

                  Gabriela Dârzan

                  DOI: 10.70744/MetaROR.418.1.rv1

                  Summary

                  This manuscript addresses an important and highly relevant topic within contemporary meta-research: inferential reproducibility in health and biomedical linear regression analyses. The study investigates whether violations of key regression assumptions contribute to differences in inferential conclusions when the same datasets are reanalysed using alternative modelling approaches.

                  The authors analysed a subset of PLOS ONE health-related papers published in 2019 that used linear regression and for which data were sufficiently available to permit computational and inferential reproducibility assessment. The manuscript combines statistical diagnostics, sensitivity analyses, bootstrapping approaches, and alternative modelling strategies to evaluate whether original conclusions remain stable under different analytical assumptions.

                  The manuscript contributes to the growing literature on reproducibility by focusing specifically on inferential reproducibility, a dimension that remains comparatively underexplored relative to computational reproducibility. The study also provides detailed practical examples of assumption checking and offers recommendations to improve transparency, promote methodological rigour, and support more reliable evidence-based decision making.

                  Overall, the manuscript is thoughtful, ambitious, and methodologically transparent. The empirical findings support the authors’ central claim that violations of statistical assumptions are frequently overlooked and may materially affect inferential precision and confidence interval estimation. However, several conceptual, methodological, and interpretational issues limit the strength of some conclusions and would benefit from clarification or refinement.

                  Major Strengths

                  1. Important and timely research question

                  The manuscript addresses a highly relevant issue within biomedical research methodology. Inferential reproducibility is increasingly recognised as a critical but understudied dimension of the broader reproducibility crisis.

                  The focus on how modelling assumptions affect inferential conclusions is valuable and practically relevant for both statistical practice and clinical interpretation.

                  2. Strong transparency and reproducibility practices

                  A major strength of the manuscript is the transparent workflow:

                  • availability of data and code;
                  • detailed explanation of diagnostic procedures;
                  • explicit reporting of sensitivity analyses;
                  • structured reproducibility framework.

                  The inclusion of GitHub materials substantially improves transparency and methodological credibility and represents good practice within Open Science.

                  3. Detailed diagnostic framework

                  The manuscript provides a comprehensive practical framework for assessing linear regression assumptions, including:

                  • linearity;
                  • independence;
                  • heteroscedasticity;
                  • outliers/influence;
                  • normality;
                  • multicollinearity.

                  The inclusion of graphical diagnostics alongside formal tests is particularly appropriate and reflects good statistical practice.

                  4. Nuanced interpretation of normality testing

                  The manuscript appropriately acknowledges limitations of formal normality tests, particularly the sensitivity of Shapiro–Wilk and Kolmogorov–Smirnov tests in large samples and their limited power in small samples.

                  The decision not to rely solely on formal significance testing, but also to incorporate graphical diagnostics and contextual judgement, is methodologically appropriate.

                  5. Constructive orientation

                  The manuscript generally avoids punitive framing and instead emphasises methodological improvement, statistical education, transparency, and context-sensitive modelling decisions. This constructive tone is appropriate for meta-research and increases the practical value of the manuscript.

                  Major Concerns

                  1. Conceptual ambiguity regarding “inferential reproducibility”

                  The manuscript would benefit from a clearer conceptual distinction between:

                  • inferential reproducibility;
                  • model robustness;
                  • sensitivity analysis;
                  • and model misspecification.

                  At several points, inferential reproducibility appears to be operationalised primarily in terms of the sensitivity of estimates to assumption violations, rather than reproducibility of substantive conclusions per se.

                  For example, some models are classified as “not inferentially reproducible” primarily because an alternative model specification yielded improved diagnostic indicators or lower AIC/BIC values, even when effect direction and substantive interpretation remained broadly similar.

                  This raises an important conceptual question:
                  Should inferential reproducibility require:

                  1. nearly identical statistical estimates;
                  2. similar inferential conclusions;
                  3. or robustness across multiple reasonable analytical approaches?

                  The manuscript would benefit from a more explicit theoretical justification for the adopted inferential reproducibility framework.

                  2. Small inferential reproducibility sample limits generalisability

                  Although the initial paper pool was substantially larger, inferential reproducibility analyses were ultimately conducted on only 14 papers and 32 models.

                  This substantially limits the generalisability of the findings and makes the reported reproducibility estimates inherently uncertain. The wide credible intervals reported by the authors appropriately reflect this uncertainty.

                  The manuscript would benefit from stronger emphasis that the study should primarily be interpreted as an exploratory empirical investigation rather than as a definitive estimate of inferential reproducibility prevalence in biomedical research more broadly.

                  3. Potential overreliance on assumption testing

                  The manuscript strongly emphasises formal assumption diagnostics. However, some discussions may unintentionally imply that violations necessarily invalidate linear regression inference.

                  Modern statistical literature generally recognises that ordinary least squares regression is generally considered relatively robust to moderate deviations from:

                  • residual normality;
                  • mild heteroscedasticity;
                  • and some forms of non-linearity;

                  particularly in moderate-to-large samples.

                  At times, the manuscript appears to treat diagnostic deviations as stronger evidence against inferential validity than may be warranted empirically.

                  The authors should better distinguish:

                  • statistical assumption violations,

                  from:

                  • practically meaningful inferential distortions.

                  This distinction is especially important because some reproduced analyses appeared to retain similar substantive conclusions despite diagnostic departures.

                  4. Subjectivity in inferential reproducibility assessment

                  An important challenge in the manuscript concerns the inherent subjectivity involved in evaluating assumption violations and selecting alternative models.

                  The manuscript appropriately incorporates expert judgement and contextual interpretation. However, different statisticians may reasonably disagree regarding:

                  • whether a violation is “minor” or “substantial”;
                  • whether a model is sufficiently robust;
                  • or whether an alternative specification is preferable.

                  This issue is particularly relevant because the inferential reproducibility classifications partly depend on judgement-based interpretation of:

                  • diagnostic plots;
                  • model fit;
                  • coefficient changes;
                  • and sensitivity analyses.

                  The manuscript would benefit from a more explicit discussion of inter-analyst variability and the possibility that different analytical teams could reach different inferential reproducibility conclusions using the same datasets.

                  The manuscript does acknowledge this issue in the limitations section. However, the point could be more explicitly connected to the interpretation of the binary inferential reproducibility classifications and to the generalisability of the study findings.

                  5. Binary classification of inferential reproducibility may oversimplify uncertainty

                  The binary Yes/No classification may oversimplify what is inherently a continuum of inferential robustness.

                  Some classifications appear difficult to interpret because:

                  • multiple criteria are combined;
                  • thresholds are partly exploratory;
                  • and judgement-based interpretation remains involved.

                  For example:

                  • 10% changes;
                  • standardised coefficient differences;
                  • CI width differences;
                  • AIC/BIC changes;
                  • and diagnostic interpretation are all incorporated simultaneously.

                  The manuscript would benefit from:

                  • a more explicit hierarchy of decision criteria,

                  or

                  • a graded reproducibility framework rather than a binary outcome classification.
                  6. Justification for inferential thresholds requires further development

                  The manuscript introduces thresholds such as:

                  • 10% coefficient changes;
                  • standardised coefficient differences of 0.1 and 0.2, to define meaningful inferential differences.

                  While the manuscript acknowledges limitations of percentage-based thresholds and appropriately explores alternative standardised metrics, the theoretical justification for selecting these particular cut-offs remains relatively limited.

                  Given that the manuscript focuses on methodological rigour and inferential interpretation, stronger justification for these thresholds would improve the conceptual robustness of the framework.

                  7. Limited contextual knowledge of original studies may influence classifications

                  The authors appropriately acknowledge that they did not contact original authors and retained original variable selections despite possible overfitting or unclear modelling rationale.

                  However, this limitation has potentially important implications.

                  Some reproduced analyses may classify studies as inferentially non-reproducible because:

                  • important contextual information;
                  • preprocessing decisions;
                  • clustering structure;
                  • transformations;
                  • or theoretical modelling rationale were unavailable.

                  This limitation should be discussed more prominently when interpreting inferential reproducibility estimates.

                  8. Risk of hindsight optimization in alternative model selection

                  In several examples, alternative models appear selected after observing diagnostic problems.

                  This introduces potential concerns regarding:

                  • researcher degrees of freedom;
                  • post hoc optimization;
                  • and analytic subjectivity.

                  Although the manuscript openly describes these procedures, stronger discussion is needed regarding:

                  • whether alternative models were pre-specified;
                  • how competing models were prioritised;
                  • and whether different analysts might reasonably arrive at different “best” models.

                  This issue is especially important given the manuscript’s focus on inferential reproducibility.

                  Minor Concerns

                  1. Clarification of terminology

                  The manuscript occasionally uses:

                  • reproducibility;
                  • robustness;
                  • sensitivity;
                  • and validity in overlapping ways.

                  A concise terminology table may improve clarity.

                  2. Interpretation of AIC/BIC improvements

                  Although the manuscript notes that AIC/BIC thresholds should be interpreted in context, some passages could more clearly distinguish improved relative model fit from stronger evidence of substantive or inferential validity. Information criteria primarily reflect relative predictive fit and penalised likelihood rather than direct evidence of model validity. This distinction should be clarified.

                  3. Reporting burden and readability

                  The manuscript is highly detailed and methodologically dense. Although this improves transparency, some sections could be shortened or reorganized for readability, particularly the extended methodological descriptions.

                  4. Discussion could better distinguish educational versus structural problems

                  The conclusion appropriately emphasises statistical education. However, inferential reproducibility problems may also reflect:

                  • publication incentives;
                  • inadequate peer review resources;
                  • insufficient methodological collaboration;
                  • and broader systemic pressures within academic publishing.

                  This broader context could be acknowledged more explicitly.

                  Suggestions for Improvement

                  1. Clarify the conceptual definition of inferential reproducibility and distinguish it more explicitly from robustness and sensitivity analysis.
                  2. Emphasize more clearly that the inferential reproducibility estimates are exploratory given the relatively small sample size.
                  3. Provide a stronger theoretical justification for the selected inferential thresholds (10%, 0.1, 0.2).
                  4. Expand discussion regarding robustness of linear regression to moderate assumption violations.
                  5. Discuss more explicitly the role of subjective judgement and potential inter-analyst variability in inferential reproducibility assessments.
                  6. Clarify how alternative models were selected and whether multiple reasonable modelling approaches could coexist.
                  7. Expand discussion regarding the limitations created by unavailable contextual information from original authors.
                  8. Consider including a summary decision flowchart explaining how inferential reproducibility classifications were determined.

                  Overall Evaluation

                  This manuscript addresses an important and insufficiently explored aspect of reproducibility research and makes a meaningful contribution to methodological discussions in biomedical research.

                  Its major strengths include:

                  • Transparency;
                  • methodological detail;
                  • practical diagnostic guidance;
                  • and focus on inferential reproducibility.

                  The empirical findings provide evidence that substantial violations of modelling assumptions are frequently insufficiently addressed in published biomedical research and may meaningfully affect inferential precision and confidence interval estimation.

                  However, the manuscript would benefit from stronger conceptual clarification regarding:

                  • inferential reproducibility;
                  • robustness;
                  • assumption violations;
                  • inferential thresholds;
                  • and model-selection subjectivity.

                  Overall, I believe the manuscript represents a valuable contribution to the meta-research literature and could represent a useful methodological contribution after clarification and refinement of several conceptual and interpretational issues.

                  I thank the authors for the transparency of the materials and analyses.

                  Competing interests: None.

                  Peer review 2

                  Konrad Hinsen

                  DOI: 10.70744/MetaROR.418.1.rv2

                  This paper examines the inferential reproducibility of health-related papers published in 2019 in the journal “PLOS One”, focusing on a basic and very widely used inference technique: linear regression. The question behind inferential reproducibility is “would someone else analyzing this data arrive at the same conclusions?” Testing for it explores mostly modeling choices made by the authors of the study being examined, which are sometimes reported and sometimes tacit. In the case of linear regression, this means testing if the conditions for applying this technique are fulfilled. The authors find that, unfortunately, this is quite often not the case.

                  The paper describes the study clearly, starting with the selection of the papers and a first check for data and code availability plus computational reproducibility, before proceeding to the core task: evaluating inferential reproducibility. This task consists of two aspects: (1) verifying the validity of applying linear regression and (2) performing linear regression a second time and compare the outcomes. Everything is documented in detail and explained with illustrative examples. This makes this paper a valuable educational resource for teaching better statistical practices to the next generation of scientists. Not being a statistician myself, I cannot judge the appropriateness of all the tests deployed, but the overall logic is clear and convincing.

                  The Discussion section points out why inferential reproducibility matters, in particular for health-related subjects that influence diagnostic and therapeutic decisions. It also emphasizes the impossibility to automate tests for inferential reproducibility, which implies that this is a task that needs to be performed by human reviewers. The authors then provide advice to such reviewers, and also to authors of papers describing statistical inference processes (not just for linear regression). This adds to the pedagogical value of the paper.

                  In their Conclusions, the authors discuss the impact of software defaults, which are often accepted uncritically, and propose approaches to improve teaching of statistical techniques to science students. They thus say, somewhat indirectly, that many scientists today have insufficient training in statistics to do their work correctly. My own experience agrees with this sentiment.

                  Overall, this paper is a welcome contribution to the improvement of statistical practices in science. It demonstrates that there are real problems, and proposes means to do better in the future. All that in a clear and comprehensible manner.

                  Finally, I have two suggestions for improving the manuscript:

                  1. Page 6 describes the HTML reports produced by the authors for each of the examined papers, but does not provide a link to them. They are in fact accessible from Ref. 19, which also contains all the code for this work. Ref. 19. is referred to in several places in the paper for different reasons. I think it would be more helpful to the reader to discuss all the material available in Ref. 19 in a single place, given that it is the central repository for the code and data backing up this study, rather than a distinct work that looks like just one cited reference out of many others.
                  2. Acronyms are mostly defined at first use, but I found a few exceptions: AIC, BIC, DFFITS, DFBETAS, COVRATIO. Also, GAM appears in the caption of Figure 1, on page 4, before being defined in the main text on page 5.

                  Competing interests: None.

                  Author response

                  DOI: 10.70744/MetaROR.418.1.ar

                  A revised version of this article is available here: https://www.medrxiv.org/content/10.64898/2026.04.07.26350296v2

                  Editorial Assessment

                  We thank the Editor for summarising the principal issues raised in the reviews. We have carefully considered each comment and revised the manuscript where additional clarification, explanation, or context was warranted.

                  We have clarified the definition and operationalisation of inferential reproducibility, its distinction from robustness and sensitivity analysis, and the roles of numerical thresholds, diagnostic assessment, information criteria, and expert judgement. We have expanded the Discussion and Limitations to provide further context regarding the sample size, inter-analyst variability, and the rationale for the binary classification, and we have also provided a flowchart to clarify the process. Supporting materials are now more clearly signposted with a direct link, and terminology and statistical acronyms have been clarified where appropriate. Detailed responses to each comment and the corresponding manuscript changes are provided below.

                  Reviewer 1: Gabriela Dârzan

                  Major concerns

                  1. Conceptual ambiguity regarding “inferential reproducibility”

                  The manuscript would benefit from a clearer conceptual distinction between:

                  • inferential reproducibility;
                  • model robustness;
                  • sensitivity analysis;
                  • and model misspecification.

                  At several points, inferential reproducibility appears to be operationalised primarily in terms of the sensitivity of estimates to assumption violations, rather than reproducibility of substantive conclusions per se.

                  For example, some models are classified as “not inferentially reproducible” primarily because an alternative model specification yielded improved diagnostic indicators or lower AIC/BIC values, even when effect direction and substantive interpretation remained broadly similar.

                  This raises an important conceptual question:
                  Should inferential reproducibility require:

                  1. nearly identical statistical estimates;
                  2. similar inferential conclusions;
                  3. or robustness across multiple reasonable analytical approaches?

                  The manuscript would benefit from a more explicit theoretical justification for the adopted inferential reproducibility framework.

                  Thank you for highlighting the need to distinguish inferential reproducibility from robustness and sensitivity analysis. We have revised the Methods to clarify that sensitivity analyses were used to assess inferential reproducibility, rather than treating these concepts as interchangeable. “The Sensitivity Analysis tab of the HTML report presents the analyses used to assess inferential reproducibility”.

                  The word “robust” has both ordinary and statistical meanings. To avoid ambiguity, we have replaced its ordinary language uses with terms such as “credible”, “trustworthy” or “stable”, depending on the context. We now use “robust” only when referring to a recognised robust statistical method. We have also included a flowchart (Figure 2) to clarify understanding of the inferential process.

                  We do not agree that inferential reproducibility was determined primarily by improved diagnostic indicators or lower AIC or BIC values. Alternative models or distributions were not fitted simply to identify the model with the most favourable information criterion. Their selection was guided by expert statistical judgement, the observed data structure, the research design, and the nature of the identified model misspecification. Information criteria were considered alongside graphical assessment, residual diagnostics, and theoretical plausibility.

                  We added to the discussion.

                  “Our framework does not require nearly identical estimates or consistency across every reasonable statistical approach. Rather, it assesses whether the original inference remains credible when the data are analysed using an alternative approach that addresses an identified problem such as an assumption violation. Although the direction of association and dichotomised statistical significance were often unchanged, these features alone do not establish inferential reproducibility. Ordinary least squares (OLS) is the standard method used to estimate coefficients in linear regression. The coefficients can still be calculated when model assumptions are not satisfied, but the validity of the estimates and conventional standard errors, confidence intervals, and p-values depends on the relevant assumptions. An unchanged direction or significance category therefore does not demonstrate that the estimated effect, its uncertainty, or any interpretation is reliable.

                  Regression assumptions are not interchangeable or additive. Several minor departures may have little practical effect, whereas a single substantial violation, such as failing to account for repeated measurements or fitting a linear model to an outcome with an incompatible mean–variance relationship, may be sufficient to undermine the estimated uncertainty and resulting inference. The final assessment was therefore expressed as a binary (Yes/No) classification of whether the original inference remained credible, rather than as a graded or additive score that could imply that a serious violation can be offset by satisfactory performance in unrelated areas.”

                  2. Small inferential reproducibility sample limits generalisability

                  Although the initial paper pool was substantially larger, inferential reproducibility analyses were ultimately conducted on only 14 papers and 32 models.

                  This substantially limits the generalisability of the findings and makes the reported reproducibility estimates inherently uncertain. The wide credible intervals reported by the authors appropriately reflect this uncertainty.

                  The manuscript would benefit from stronger emphasis that the study should primarily be interpreted as an exploratory empirical investigation rather than as a definitive estimate of inferential reproducibility prevalence in biomedical research more broadly.

                  We agree that the relatively small number of papers available for inferential reproducibility assessment limits the precision of the estimated proportion. The study was designed as a pilot because there was insufficient empirical evidence to support a formal sample-size calculation. To our knowledge, few studies have estimated inferential reproducibility using a random sample of published research, so these findings provide useful preliminary evidence despite the limited sample size.

                  The manuscript already identifies the study as exploratory and explicitly cautions readers about the small number of papers and the resulting uncertainty. The results state:

                  “Given the small number of papers (N = 14) and the resulting uncertainty reflected in the wide credible interval, the estimate should be interpreted cautiously.”

                  We edited the limitations to acknowledge the small sample size.

                  “Our results should be interpreted with caution, as we only assessed inferential reproducibility in papers that were computationally reproducible, a subset that may disproportionately represent methodologically skilled authors and was conducted using a relatively small number of papers.”

                  3. Potential overreliance on assumption testing

                  The manuscript strongly emphasises formal assumption diagnostics. However, some discussions may unintentionally imply that violations necessarily invalidate linear regression inference.

                  Modern statistical literature generally recognises that ordinary least squares regression is generally considered relatively robust to moderate deviations from:

                  • residual normality;
                  • mild heteroscedasticity;
                  • and some forms of non-linearity;

                  particularly in moderate-to-large samples.

                  At times, the manuscript appears to treat diagnostic deviations as stronger evidence against inferential validity than may be warranted empirically.

                  The authors should better distinguish:

                  • statistical assumption violations,

                  from:

                  • practically meaningful inferential distortions.

                  This distinction is especially important because some reproduced analyses appeared to retain similar substantive conclusions despite diagnostic departures.

                  We respectfully disagree that the manuscript treats assumption violations as automatically invalidating inference. The original manuscript explicitly stated that minor departures were considered acceptable when they were unlikely to materially affect estimates, confidence intervals, model adequacy, or interpretation. It also emphasises graphical assessment and contextual statistical judgement rather than reliance on formal tests alone.

                  Although we applied a comprehensive set of diagnostic tests and plots, this was intended to provide a complete assessment of model behaviour and to examine how often identified departures had little practical consequence. Some formal tests, including tests of residual normality, would not usually be used by statisticians as a basis for judging whether an assumption has been violated, because their results are highly dependent on sample size. They were therefore interpreted alongside graphical diagnostics, the magnitude and pattern of departures, and their observed effect on the resulting inference. Table 2 describes the assumption violations in each paper and then puts their severity into context.

                  4. Subjectivity in inferential reproducibility assessment

                  An important challenge in the manuscript concerns the inherent subjectivity involved in evaluating assumption violations and selecting alternative models.

                  The manuscript appropriately incorporates expert judgement and contextual interpretation. However, different statisticians may reasonably disagree regarding:

                  • whether a violation is “minor” or “substantial”;
                  • whether a model is sufficiently robust;
                  • or whether an alternative specification is preferable.

                  This issue is particularly relevant because the inferential reproducibility classifications partly depend on judgement-based interpretation of:

                  • diagnostic plots;
                  • model fit;
                  • coefficient changes;
                  • and sensitivity analyses.

                  The manuscript would benefit from a more explicit discussion of inter-analyst variability and the possibility that different analytical teams could reach different inferential reproducibility conclusions using the same datasets.

                  The manuscript does acknowledge this issue in the limitations section. However, the point could be more explicitly connected to the interpretation of the binary inferential reproducibility classifications and to the generalisability of the study findings.

                  We agree that different statisticians may reasonably reach different conclusions when assessing model assumptions and selecting alternative analyses, and this was explicitly acknowledged in the limitations. However, such inter-analyst variability is inherent to statistical modelling and is directly relevant to the concept of inferential reproducibility examined in this study, rather than being introduced solely by our framework. The original discussion considers evidence from multi-analyst studies showing that, even among experienced analysts examining the same data, different defensible analytical choices and conclusions may arise.

                  We added to the limitations

                  “Assumption violations were not considered consequential solely because they were present, as their effects depend on their nature, severity, and the context of the analysis. When a violation was identified, expert judgement was used to assess its potential importance, and sensitivity analyses were conducted to determine whether addressing it materially affected the estimates or resulting inferences. This process required judgement in selecting appropriate alternative models and determining whether observed differences were substantively meaningful.”

                  We do not consider the binary classification of inferential reproducibility a limitation. In statistical practice, statisticians must ultimately determine whether a model provides an adequate basis for the intended inference. The binary inferential reproducibility (Yes/No) classification summarised this overall judgement by indicating whether the original inference remained credible after an assumption violation had been addressed.

                  The original manuscript provides detailed paper-level reports, referenced in Table 2, describing the diagnostic findings, whether departures were considered minor or substantial, the rationale for each sensitivity analysis, and the basis for the final classification. The original HTML files are also available on GitHub, allowing readers to examine the evidence underlying each decision and form their own judgement.

                  5. Binary classification of inferential reproducibility may oversimplify uncertainty

                  The binary Yes/No classification may oversimplify what is inherently a continuum of inferential robustness.

                  Some classifications appear difficult to interpret because:

                  • multiple criteria are combined;
                  • thresholds are partly exploratory;
                  • and judgement-based interpretation remains involved.

                  For example:

                  • 10% changes;
                  • standardised coefficient differences;
                  • CI width differences;
                  • AIC/BIC changes;
                  • and diagnostic interpretation are all incorporated simultaneously.

                  The manuscript would benefit from:

                  • a more explicit hierarchy of decision criteria,

                  or

                  • a graded reproducibility framework rather than a binary outcome classification.

                  Inferential reproducibility is inherently complex, and the framework was developed to reflect the reasoning used in statistical practice. Numerical thresholds were used to support consistent assessment of changes in estimates and precision, but they were interpreted alongside study design, model diagnostics, the nature and severity of any assumption violations, and the suitability of the alternative analysis. The final Yes/No classification therefore represented an overall statistical judgement about whether the original inference remained credible, rather than being determined by any single threshold or a fixed combination of criteria.

                  6. Justification for inferential thresholds requires further development

                  The manuscript introduces thresholds such as:

                  • 10% coefficient changes;
                  • standardised coefficient differences of 0.1 and 0.2, to define meaningful inferential differences.

                  While the manuscript acknowledges limitations of percentage-based thresholds and appropriately explores alternative standardised metrics, the theoretical justification for selecting these particular cut-offs remains relatively limited.

                  Given that the manuscript focuses on methodological rigour and inferential interpretation, stronger justification for these thresholds would improve the conceptual robustness of the framework.

                  The rationale and limitations of these thresholds were described in the manuscript, including the instability of percentage changes when coefficients are close to zero and the use of standardised differences as an alternative metric. The thresholds were intended as pragmatic decision aids rather than universally validated cut-offs and were interpreted alongside diagnostic findings and expert judgement.

                  We have added the following clarification to the limitations:

                  “These pragmatic cut-offs were informed by the research team’s experience in judging whether small differences were substantively meaningful and by conventional benchmarks for small effect sizes. However, further research is needed to validate these thresholds.”

                  7. Limited contextual knowledge of original studies may influence classifications

                  The authors appropriately acknowledge that they did not contact original authors and retained original variable selections despite possible overfitting or unclear modelling rationale.

                  However, this limitation has potentially important implications.

                  Some reproduced analyses may classify studies as inferentially non-reproducible because:

                  • important contextual information;
                  • preprocessing decisions;
                  • clustering structure;
                  • transformations;
                  • or theoretical modelling rationale were unavailable.

                  This limitation should be discussed more prominently when interpreting inferential reproducibility estimates.

                  Thank you for this comment. The manuscript already acknowledges that unavailable contextual information may have affected some assessments. However, inferential reproducibility requires sufficient reporting and documentation for an independent analyst to understand whether the model and its assumptions were appropriate. If preprocessing decisions, transformations, clustering structures or modelling rationales cannot be recovered from the article and shared materials, this is directly relevant to inferential reproducibility.

                  We have added to the limitations

                  “Accordingly, the individual classifications and overall inferential reproducibility estimate reflect what could be independently assessed from the published articles and shared materials; additional undocumented contextual information may have altered some assessments.”

                  8. Risk of hindsight optimization in alternative model selection

                  In several examples, alternative models appear selected after observing diagnostic problems.

                  This introduces potential concerns regarding:

                  • researcher degrees of freedom;
                  • post hoc optimization;
                  • and analytic subjectivity.

                  Although the manuscript openly describes these procedures, stronger discussion is needed regarding:

                  • whether alternative models were pre-specified;
                  • how competing models were prioritised;
                  • and whether different analysts might reasonably arrive at different “best” models.

                  This issue is especially important given the manuscript’s focus on inferential reproducibility.

                  We do not consider these analyses to constitute post hoc optimisation because they were selected to address specific diagnostic concerns rather than to obtain more favourable results. Although evaluating alternative models after examining the data may increase the risk of overfitting or producing findings that are specific to the observed dataset, this risk was reduced by considering graphical diagnostics, model fit, the suitability of each model for the data, and the consistency of findings across other reasonable analyses.

                  The exact alternative model cannot be prespecified because the relevant form of misspecification may become evident only after the original model and data have been examined. The study therefore prespecified a structured assessment framework rather than a fixed alternative model for every possible problem. Additional analyses were undertaken only when supported by the diagnostics and were selected to address the specific concern identified.

                  The aim was not to identify a single “best” model, but to determine whether the original inference remained credible under a reasonable analysis that addressed the identified problem. The rationale for each sensitivity analysis is documented in the paper-level reports (Table 2), and the manuscript acknowledges that another statistician might sometimes select a different, but equally defensible, approach.

                  Minor Concerns

                  1. Clarification of terminology

                  The manuscript occasionally uses:

                  • reproducibility;
                  • robustness;
                  • sensitivity;
                  • and validity in overlapping ways.

                  A concise terminology table may improve clarity.

                  Changes were made to improve clarity. For further detail, see major concern 1.

                  2. Interpretation of AIC/BIC improvements

                  Although the manuscript notes that AIC/BIC thresholds should be interpreted in context, some passages could more clearly distinguish improved relative model fit from stronger evidence of substantive or inferential validity. Information criteria primarily reflect relative predictive fit and penalised likelihood rather than direct evidence of model validity. This distinction should be clarified.

                  We agree that AIC and BIC assess relative support among candidate models and do not independently establish substantive or inferential validity. In the manuscript, they were used alongside graphical assessment, diagnostic findings, theoretical plausibility, model appropriateness, and the resulting estimates. We have nevertheless added a paragraph to the Discussion, following the discussion of influential observations, to clarify why information criteria should not be interpreted in isolation.

                  “Papers 32 and 46 also illustrate why models should not be selected solely on the basis of AIC or BIC (Table 2). Although these criteria compare the relative fit of candidate models while penalising model complexity, the model with the lowest value is not necessarily the most scientifically plausible. In these examples, polynomial terms initially appeared to improve model fit, but graphical assessment showed that the apparent relationships were strongly influenced by extreme observations. A higher-order polynomial may similarly improve AIC or BIC by closely following features specific to the observed dataset while representing a relationship that is difficult to justify or unlikely to generalise. Information criteria should therefore be interpreted alongside graphical assessment, residual diagnostics, theoretical plausibility, and the influence of individual observations.”

                  3. Reporting burden and readability

                  The manuscript is highly detailed and methodologically dense. Although this improves transparency, some sections could be shortened or reorganized for readability, particularly the extended methodological descriptions.

                  We appreciate the reviewer’s concern regarding readability. However, we do not agree that the methodological descriptions should be substantially shortened. Assessing inferential reproducibility requires several analytical decisions that are not yet governed by a single established framework, including assessing model assumptions, selecting targeted sensitivity analyses, and determining whether results from alternative analyses can be validly compared. Sufficient methodological detail is therefore necessary to make the statistical reasoning transparent and reproducible. To support readability, we have organised the methodological material using clear headings and subheadings, allowing readers to navigate directly to the sections most relevant to them.

                  4. Discussion could better distinguish educational versus structural problems

                  The conclusion appropriately emphasises statistical education. However, inferential reproducibility problems may also reflect:

                  • publication incentives;
                  • inadequate peer review resources;
                  • insufficient methodological collaboration;
                  • and broader systemic pressures within academic publishing.

                  This broader context could be acknowledged more explicitly.

                  Thank you for your comment; we have added a paragraph discussing structural problems in the conclusion.

                  “However, inferential reproducibility problems should not be attributed solely to gaps in individual researchers’ statistical knowledge. Publication incentives, including pressure to publish, and broader systemic pressures may limit the time, resources, and support available for careful model development, diagnostic assessment, and transparent reporting [73]. Limited peer-review capacity and inconsistent access to methodological expertise may further reduce opportunities for statistical problems to be identified before publication [74, 75].”

                  Suggestions for Improvement

                  1. Clarify the conceptual definition of inferential reproducibility and distinguish it more explicitly from robustness and sensitivity analysis.

                  Changes were made to improve clarity. For further detail, see major concern 1.

                  2. Emphasize more clearly that the inferential reproducibility estimates are exploratory given the relatively small sample size.

                  The small sample size was reiterated in the limitations section. For details, see major concern 2.

                  3. Provide a stronger theoretical justification for the selected inferential thresholds (10%, 0.1, 0.2).

                  We have added clarification to the limitations. For further detail, see major concern 6.

                  4. Expand discussion regarding robustness of linear regression to moderate assumption violations.

                  We have added paragraphs discussing this in both the discussion and limitations; for more detail, see major concern 1 and 4.

                  5. Discuss more explicitly the role of subjective judgement and potential inter-analyst variability in inferential reproducibility assessments.

                  We added a discussion point to the limitations. For further detail, see major concern 4.

                  6. Clarify how alternative models were selected and whether multiple reasonable modelling approaches could coexist.

                  We added further detail to clarify how alternative models were selected. For further detail, see major concern 1.

                  7. Expand discussion regarding the limitations created by unavailable contextual information from original authors.

                  We have added a point to the limitations section. For further detail, see major concern 7.

                  8. Consider including a summary decision flowchart explaining how inferential reproducibility classifications were determined.

                  We have included a flowchart (Figure 2) to clarify understanding of the inferential reproducibility process.

                  Overall Evaluation

                  Thank you for your review. We have clarified the methodological framework and the distinction between inferential reproducibility, robustness, and sensitivity analysis. We have also expanded the explanation of how assumption violations, numerical thresholds, information criteria, and expert judgement informed the final classifications. Additional revisions address the interpretation of the binary classification, model-selection subjectivity, inter-analyst variability, and the study’s limitations. Detailed responses to each point are provided above.

                  Reviewer 2:  Konrad Hinsen

                  1. Page 6 describes the HTML reports produced by the authors for each of the examined papers, but does not provide a link to them. They are in fact accessible from Ref. 19, which also contains all the code for this work. Ref. 19. is referred to in several places in the paper for different reasons. I think it would be more helpful to the reader to discuss all the material available in Ref. 19 in a single place, given that it is the central repository for the code and data backing up this study, rather than a distinct work that looks like just one cited reference out of many others.

                  While not directly in the manuscript, the data availability statement is available on medRxiv and has not been duplicated, as this metadata is usually included on the paper’s first page when it is published, along with conflict of interest, etc.

                  Data Availability

                  The data and a reproducible R Quarto file used to produce this paper, including tables, figures, and code, have been stored in a GitHub repository and can be cited using the Zenodo DOI.

                  https://github.com/Lee-V-Jones/Reproducibility

                  https://doi.org/10.5281/zenodo.19448969

                  However, for ease of access, we have added a direct link in the text to the 95 HTML reports. “The reports are publicly available at https://github.com/Lee-V-Jones/Reproducibility/blob/main/README.md

                  2. Acronyms are mostly defined at first use, but I found a few exceptions: AIC, BIC, DFFITS, DFBETAS, COVRATIO. Also, GAM appears in the caption of Figure 1, on page 4, before being defined in the main text on page 5.

                  Thanks for picking this up. This was an oversight: we defined acronyms for GAM, AIC, BIC, ANOVA, DFFITS, DFBETAS, and COVRATIO the first time they were used. We have edited Figure 1 to define Generalised Additive Models (GAM).

                  Leave a comment