Published at MetaROR
August 10, 2026
Table of contents
An empirical assessment of inferential reproducibility of linear regression in health and biomedical research papers
1 Research Methods Group, Faculty of Health, School of Public Health and Social Work, Queensland University of Technology, Kelvin Grove, Queensland, Australia,
2 AusHSI, Centre for Healthcare Transformation, Faculty of Health, School of Public Health and Social Work, Queensland University of Technology, Kelvin Grove, Queensland, Australia,
3 Statistics Unit, QIMR Berghofer Medical Research Institute, Herston, Queensland, Australia,
Originally published on April 7, 2026 at:
Editors
Kathryn Zeiler
Wolfgang Kaltenbrunner
Editorial assessment
by Wolfgang Kaltenbrunner
The manuscript presents an empirical assessment of the inferential reproducibility of linear regression analyses in health and biomedical research. Both reviewers agree that it addresses a timely and important question within reproducibility research and highlight the study’s methodological transparency, detailed diagnostic framework, and strong pedagogical value. They note that the paper makes a valuable contribution by moving beyond computational reproducibility to examine how violations of statistical assumptions may affect inferential conclusions and by providing practical guidance for researchers, reviewers, and students.
At the same time, both reviews identify areas where the manuscript could be strengthened. A recurring theme concerns the conceptual framing of inferential reproducibility, including its distinction from robustness and sensitivity analysis, the justification of inferential thresholds, and the role of analyst judgement in assessing reproducibility. Reviewers also encourage a more explicit discussion of the exploratory nature of the findings given the relatively small inferential reproducibility sample, the robustness of linear regression to moderate assumption violations, and the subjectivity involved in selecting and evaluating alternative models. Additional suggestions include improving access to supporting materials, clarifying terminology and statistical acronyms, and further contextualizing the study’s findings and limitations. Overall, the reviewers and I agree that this is a valuable and constructive contribution to the meta-research literature that would benefit from clarification and refinement of several conceptual and interpretive aspects.
Peer review 1
Summary
This manuscript addresses an important and highly relevant topic within contemporary meta-research: inferential reproducibility in health and biomedical linear regression analyses. The study investigates whether violations of key regression assumptions contribute to differences in inferential conclusions when the same datasets are reanalysed using alternative modelling approaches.
The authors analysed a subset of PLOS ONE health-related papers published in 2019 that used linear regression and for which data were sufficiently available to permit computational and inferential reproducibility assessment. The manuscript combines statistical diagnostics, sensitivity analyses, bootstrapping approaches, and alternative modelling strategies to evaluate whether original conclusions remain stable under different analytical assumptions.
The manuscript contributes to the growing literature on reproducibility by focusing specifically on inferential reproducibility, a dimension that remains comparatively underexplored relative to computational reproducibility. The study also provides detailed practical examples of assumption checking and offers recommendations to improve transparency, promote methodological rigour, and support more reliable evidence-based decision making.
Overall, the manuscript is thoughtful, ambitious, and methodologically transparent. The empirical findings support the authors’ central claim that violations of statistical assumptions are frequently overlooked and may materially affect inferential precision and confidence interval estimation. However, several conceptual, methodological, and interpretational issues limit the strength of some conclusions and would benefit from clarification or refinement.
Major Strengths
1. Important and timely research question
The manuscript addresses a highly relevant issue within biomedical research methodology. Inferential reproducibility is increasingly recognised as a critical but understudied dimension of the broader reproducibility crisis.
The focus on how modelling assumptions affect inferential conclusions is valuable and practically relevant for both statistical practice and clinical interpretation.
2. Strong transparency and reproducibility practices
A major strength of the manuscript is the transparent workflow:
- availability of data and code;
- detailed explanation of diagnostic procedures;
- explicit reporting of sensitivity analyses;
- structured reproducibility framework.
The inclusion of GitHub materials substantially improves transparency and methodological credibility and represents good practice within Open Science.
3. Detailed diagnostic framework
The manuscript provides a comprehensive practical framework for assessing linear regression assumptions, including:
- linearity;
- independence;
- heteroscedasticity;
- outliers/influence;
- normality;
- multicollinearity.
The inclusion of graphical diagnostics alongside formal tests is particularly appropriate and reflects good statistical practice.
4. Nuanced interpretation of normality testing
The manuscript appropriately acknowledges limitations of formal normality tests, particularly the sensitivity of Shapiro–Wilk and Kolmogorov–Smirnov tests in large samples and their limited power in small samples.
The decision not to rely solely on formal significance testing, but also to incorporate graphical diagnostics and contextual judgement, is methodologically appropriate.
5. Constructive orientation
The manuscript generally avoids punitive framing and instead emphasises methodological improvement, statistical education, transparency, and context-sensitive modelling decisions. This constructive tone is appropriate for meta-research and increases the practical value of the manuscript.
Major Concerns
1. Conceptual ambiguity regarding “inferential reproducibility”
The manuscript would benefit from a clearer conceptual distinction between:
- inferential reproducibility;
- model robustness;
- sensitivity analysis;
- and model misspecification.
At several points, inferential reproducibility appears to be operationalised primarily in terms of the sensitivity of estimates to assumption violations, rather than reproducibility of substantive conclusions per se.
For example, some models are classified as “not inferentially reproducible” primarily because an alternative model specification yielded improved diagnostic indicators or lower AIC/BIC values, even when effect direction and substantive interpretation remained broadly similar.
This raises an important conceptual question:
Should inferential reproducibility require:
- nearly identical statistical estimates;
- similar inferential conclusions;
- or robustness across multiple reasonable analytical approaches?
The manuscript would benefit from a more explicit theoretical justification for the adopted inferential reproducibility framework.
2. Small inferential reproducibility sample limits generalisability
Although the initial paper pool was substantially larger, inferential reproducibility analyses were ultimately conducted on only 14 papers and 32 models.
This substantially limits the generalisability of the findings and makes the reported reproducibility estimates inherently uncertain. The wide credible intervals reported by the authors appropriately reflect this uncertainty.
The manuscript would benefit from stronger emphasis that the study should primarily be interpreted as an exploratory empirical investigation rather than as a definitive estimate of inferential reproducibility prevalence in biomedical research more broadly.
3. Potential overreliance on assumption testing
The manuscript strongly emphasises formal assumption diagnostics. However, some discussions may unintentionally imply that violations necessarily invalidate linear regression inference.
Modern statistical literature generally recognises that ordinary least squares regression is generally considered relatively robust to moderate deviations from:
- residual normality;
- mild heteroscedasticity;
- and some forms of non-linearity;
particularly in moderate-to-large samples.
At times, the manuscript appears to treat diagnostic deviations as stronger evidence against inferential validity than may be warranted empirically.
The authors should better distinguish:
- statistical assumption violations,
from:
- practically meaningful inferential distortions.
This distinction is especially important because some reproduced analyses appeared to retain similar substantive conclusions despite diagnostic departures.
4. Subjectivity in inferential reproducibility assessment
An important challenge in the manuscript concerns the inherent subjectivity involved in evaluating assumption violations and selecting alternative models.
The manuscript appropriately incorporates expert judgement and contextual interpretation. However, different statisticians may reasonably disagree regarding:
- whether a violation is “minor” or “substantial”;
- whether a model is sufficiently robust;
- or whether an alternative specification is preferable.
This issue is particularly relevant because the inferential reproducibility classifications partly depend on judgement-based interpretation of:
- diagnostic plots;
- model fit;
- coefficient changes;
- and sensitivity analyses.
The manuscript would benefit from a more explicit discussion of inter-analyst variability and the possibility that different analytical teams could reach different inferential reproducibility conclusions using the same datasets.
The manuscript does acknowledge this issue in the limitations section. However, the point could be more explicitly connected to the interpretation of the binary inferential reproducibility classifications and to the generalisability of the study findings.
5. Binary classification of inferential reproducibility may oversimplify uncertainty
The binary Yes/No classification may oversimplify what is inherently a continuum of inferential robustness.
Some classifications appear difficult to interpret because:
- multiple criteria are combined;
- thresholds are partly exploratory;
- and judgement-based interpretation remains involved.
For example:
- 10% changes;
- standardised coefficient differences;
- CI width differences;
- AIC/BIC changes;
- and diagnostic interpretation are all incorporated simultaneously.
The manuscript would benefit from:
- a more explicit hierarchy of decision criteria,
or
- a graded reproducibility framework rather than a binary outcome classification.
6. Justification for inferential thresholds requires further development
The manuscript introduces thresholds such as:
- 10% coefficient changes;
- standardised coefficient differences of 0.1 and 0.2, to define meaningful inferential differences.
While the manuscript acknowledges limitations of percentage-based thresholds and appropriately explores alternative standardised metrics, the theoretical justification for selecting these particular cut-offs remains relatively limited.
Given that the manuscript focuses on methodological rigour and inferential interpretation, stronger justification for these thresholds would improve the conceptual robustness of the framework.
7. Limited contextual knowledge of original studies may influence classifications
The authors appropriately acknowledge that they did not contact original authors and retained original variable selections despite possible overfitting or unclear modelling rationale.
However, this limitation has potentially important implications.
Some reproduced analyses may classify studies as inferentially non-reproducible because:
- important contextual information;
- preprocessing decisions;
- clustering structure;
- transformations;
- or theoretical modelling rationale were unavailable.
This limitation should be discussed more prominently when interpreting inferential reproducibility estimates.
8. Risk of hindsight optimization in alternative model selection
In several examples, alternative models appear selected after observing diagnostic problems.
This introduces potential concerns regarding:
- researcher degrees of freedom;
- post hoc optimization;
- and analytic subjectivity.
Although the manuscript openly describes these procedures, stronger discussion is needed regarding:
- whether alternative models were pre-specified;
- how competing models were prioritised;
- and whether different analysts might reasonably arrive at different “best” models.
This issue is especially important given the manuscript’s focus on inferential reproducibility.
Minor Concerns
1. Clarification of terminology
The manuscript occasionally uses:
- reproducibility;
- robustness;
- sensitivity;
- and validity in overlapping ways.
A concise terminology table may improve clarity.
2. Interpretation of AIC/BIC improvements
Although the manuscript notes that AIC/BIC thresholds should be interpreted in context, some passages could more clearly distinguish improved relative model fit from stronger evidence of substantive or inferential validity. Information criteria primarily reflect relative predictive fit and penalised likelihood rather than direct evidence of model validity. This distinction should be clarified.
3. Reporting burden and readability
The manuscript is highly detailed and methodologically dense. Although this improves transparency, some sections could be shortened or reorganized for readability, particularly the extended methodological descriptions.
4. Discussion could better distinguish educational versus structural problems
The conclusion appropriately emphasises statistical education. However, inferential reproducibility problems may also reflect:
- publication incentives;
- inadequate peer review resources;
- insufficient methodological collaboration;
- and broader systemic pressures within academic publishing.
This broader context could be acknowledged more explicitly.
Suggestions for Improvement
- Clarify the conceptual definition of inferential reproducibility and distinguish it more explicitly from robustness and sensitivity analysis.
- Emphasize more clearly that the inferential reproducibility estimates are exploratory given the relatively small sample size.
- Provide a stronger theoretical justification for the selected inferential thresholds (10%, 0.1, 0.2).
- Expand discussion regarding robustness of linear regression to moderate assumption violations.
- Discuss more explicitly the role of subjective judgement and potential inter-analyst variability in inferential reproducibility assessments.
- Clarify how alternative models were selected and whether multiple reasonable modelling approaches could coexist.
- Expand discussion regarding the limitations created by unavailable contextual information from original authors.
- Consider including a summary decision flowchart explaining how inferential reproducibility classifications were determined.
Overall Evaluation
This manuscript addresses an important and insufficiently explored aspect of reproducibility research and makes a meaningful contribution to methodological discussions in biomedical research.
Its major strengths include:
- Transparency;
- methodological detail;
- practical diagnostic guidance;
- and focus on inferential reproducibility.
The empirical findings provide evidence that substantial violations of modelling assumptions are frequently insufficiently addressed in published biomedical research and may meaningfully affect inferential precision and confidence interval estimation.
However, the manuscript would benefit from stronger conceptual clarification regarding:
- inferential reproducibility;
- robustness;
- assumption violations;
- inferential thresholds;
- and model-selection subjectivity.
Overall, I believe the manuscript represents a valuable contribution to the meta-research literature and could represent a useful methodological contribution after clarification and refinement of several conceptual and interpretational issues.
I thank the authors for the transparency of the materials and analyses.
Peer review 2
This paper examines the inferential reproducibility of health-related papers published in 2019 in the journal “PLOS One”, focusing on a basic and very widely used inference technique: linear regression. The question behind inferential reproducibility is “would someone else analyzing this data arrive at the same conclusions?” Testing for it explores mostly modeling choices made by the authors of the study being examined, which are sometimes reported and sometimes tacit. In the case of linear regression, this means testing if the conditions for applying this technique are fulfilled. The authors find that, unfortunately, this is quite often not the case.
The paper describes the study clearly, starting with the selection of the papers and a first check for data and code availability plus computational reproducibility, before proceeding to the core task: evaluating inferential reproducibility. This task consists of two aspects: (1) verifying the validity of applying linear regression and (2) performing linear regression a second time and compare the outcomes. Everything is documented in detail and explained with illustrative examples. This makes this paper a valuable educational resource for teaching better statistical practices to the next generation of scientists. Not being a statistician myself, I cannot judge the appropriateness of all the tests deployed, but the overall logic is clear and convincing.
The Discussion section points out why inferential reproducibility matters, in particular for health-related subjects that influence diagnostic and therapeutic decisions. It also emphasizes the impossibility to automate tests for inferential reproducibility, which implies that this is a task that needs to be performed by human reviewers. The authors then provide advice to such reviewers, and also to authors of papers describing statistical inference processes (not just for linear regression). This adds to the pedagogical value of the paper.
In their Conclusions, the authors discuss the impact of software defaults, which are often accepted uncritically, and propose approaches to improve teaching of statistical techniques to science students. They thus say, somewhat indirectly, that many scientists today have insufficient training in statistics to do their work correctly. My own experience agrees with this sentiment.
Overall, this paper is a welcome contribution to the improvement of statistical practices in science. It demonstrates that there are real problems, and proposes means to do better in the future. All that in a clear and comprehensible manner.
Finally, I have two suggestions for improving the manuscript:
- Page 6 describes the HTML reports produced by the authors for each of the examined papers, but does not provide a link to them. They are in fact accessible from Ref. 19, which also contains all the code for this work. Ref. 19. is referred to in several places in the paper for different reasons. I think it would be more helpful to the reader to discuss all the material available in Ref. 19 in a single place, given that it is the central repository for the code and data backing up this study, rather than a distinct work that looks like just one cited reference out of many others.
- Acronyms are mostly defined at first use, but I found a few exceptions: AIC, BIC, DFFITS, DFBETAS, COVRATIO. Also, GAM appears in the caption of Figure 1, on page 4, before being defined in the main text on page 5.
Author response
A revised version of this article is available here: https://www.medrxiv.org/content/10.64898/2026.04.07.26350296v2
Editorial Assessment
We thank the Editor for summarising the principal issues raised in the reviews. We have carefully considered each comment and revised the manuscript where additional clarification, explanation, or context was warranted.
We have clarified the definition and operationalisation of inferential reproducibility, its distinction from robustness and sensitivity analysis, and the roles of numerical thresholds, diagnostic assessment, information criteria, and expert judgement. We have expanded the Discussion and Limitations to provide further context regarding the sample size, inter-analyst variability, and the rationale for the binary classification, and we have also provided a flowchart to clarify the process. Supporting materials are now more clearly signposted with a direct link, and terminology and statistical acronyms have been clarified where appropriate. Detailed responses to each comment and the corresponding manuscript changes are provided below.
Reviewer 1: Gabriela Dârzan
Major concerns
1. Conceptual ambiguity regarding “inferential reproducibility”
The manuscript would benefit from a clearer conceptual distinction between:
- inferential reproducibility;
- model robustness;
- sensitivity analysis;
- and model misspecification.
At several points, inferential reproducibility appears to be operationalised primarily in terms of the sensitivity of estimates to assumption violations, rather than reproducibility of substantive conclusions per se.
For example, some models are classified as “not inferentially reproducible” primarily because an alternative model specification yielded improved diagnostic indicators or lower AIC/BIC values, even when effect direction and substantive interpretation remained broadly similar.
This raises an important conceptual question:
Should inferential reproducibility require:
- nearly identical statistical estimates;
- similar inferential conclusions;
- or robustness across multiple reasonable analytical approaches?
The manuscript would benefit from a more explicit theoretical justification for the adopted inferential reproducibility framework.
Thank you for highlighting the need to distinguish inferential reproducibility from robustness and sensitivity analysis. We have revised the Methods to clarify that sensitivity analyses were used to assess inferential reproducibility, rather than treating these concepts as interchangeable. “The Sensitivity Analysis tab of the HTML report presents the analyses used to assess inferential reproducibility”.
The word “robust” has both ordinary and statistical meanings. To avoid ambiguity, we have replaced its ordinary language uses with terms such as “credible”, “trustworthy” or “stable”, depending on the context. We now use “robust” only when referring to a recognised robust statistical method. We have also included a flowchart (Figure 2) to clarify understanding of the inferential process.
We do not agree that inferential reproducibility was determined primarily by improved diagnostic indicators or lower AIC or BIC values. Alternative models or distributions were not fitted simply to identify the model with the most favourable information criterion. Their selection was guided by expert statistical judgement, the observed data structure, the research design, and the nature of the identified model misspecification. Information criteria were considered alongside graphical assessment, residual diagnostics, and theoretical plausibility.
We added to the discussion.
“Our framework does not require nearly identical estimates or consistency across every reasonable statistical approach. Rather, it assesses whether the original inference remains credible when the data are analysed using an alternative approach that addresses an identified problem such as an assumption violation. Although the direction of association and dichotomised statistical significance were often unchanged, these features alone do not establish inferential reproducibility. Ordinary least squares (OLS) is the standard method used to estimate coefficients in linear regression. The coefficients can still be calculated when model assumptions are not satisfied, but the validity of the estimates and conventional standard errors, confidence intervals, and p-values depends on the relevant assumptions. An unchanged direction or significance category therefore does not demonstrate that the estimated effect, its uncertainty, or any interpretation is reliable.
Regression assumptions are not interchangeable or additive. Several minor departures may have little practical effect, whereas a single substantial violation, such as failing to account for repeated measurements or fitting a linear model to an outcome with an incompatible mean–variance relationship, may be sufficient to undermine the estimated uncertainty and resulting inference. The final assessment was therefore expressed as a binary (Yes/No) classification of whether the original inference remained credible, rather than as a graded or additive score that could imply that a serious violation can be offset by satisfactory performance in unrelated areas.”
2. Small inferential reproducibility sample limits generalisability
Although the initial paper pool was substantially larger, inferential reproducibility analyses were ultimately conducted on only 14 papers and 32 models.
This substantially limits the generalisability of the findings and makes the reported reproducibility estimates inherently uncertain. The wide credible intervals reported by the authors appropriately reflect this uncertainty.
The manuscript would benefit from stronger emphasis that the study should primarily be interpreted as an exploratory empirical investigation rather than as a definitive estimate of inferential reproducibility prevalence in biomedical research more broadly.
We agree that the relatively small number of papers available for inferential reproducibility assessment limits the precision of the estimated proportion. The study was designed as a pilot because there was insufficient empirical evidence to support a formal sample-size calculation. To our knowledge, few studies have estimated inferential reproducibility using a random sample of published research, so these findings provide useful preliminary evidence despite the limited sample size.
The manuscript already identifies the study as exploratory and explicitly cautions readers about the small number of papers and the resulting uncertainty. The results state:
“Given the small number of papers (N = 14) and the resulting uncertainty reflected in the wide credible interval, the estimate should be interpreted cautiously.”
We edited the limitations to acknowledge the small sample size.
“Our results should be interpreted with caution, as we only assessed inferential reproducibility in papers that were computationally reproducible, a subset that may disproportionately represent methodologically skilled authors and was conducted using a relatively small number of papers.”
3. Potential overreliance on assumption testing
The manuscript strongly emphasises formal assumption diagnostics. However, some discussions may unintentionally imply that violations necessarily invalidate linear regression inference.
Modern statistical literature generally recognises that ordinary least squares regression is generally considered relatively robust to moderate deviations from:
- residual normality;
- mild heteroscedasticity;
- and some forms of non-linearity;
particularly in moderate-to-large samples.
At times, the manuscript appears to treat diagnostic deviations as stronger evidence against inferential validity than may be warranted empirically.
The authors should better distinguish:
- statistical assumption violations,
from:
- practically meaningful inferential distortions.
This distinction is especially important because some reproduced analyses appeared to retain similar substantive conclusions despite diagnostic departures.
We respectfully disagree that the manuscript treats assumption violations as automatically invalidating inference. The original manuscript explicitly stated that minor departures were considered acceptable when they were unlikely to materially affect estimates, confidence intervals, model adequacy, or interpretation. It also emphasises graphical assessment and contextual statistical judgement rather than reliance on formal tests alone.
Although we applied a comprehensive set of diagnostic tests and plots, this was intended to provide a complete assessment of model behaviour and to examine how often identified departures had little practical consequence. Some formal tests, including tests of residual normality, would not usually be used by statisticians as a basis for judging whether an assumption has been violated, because their results are highly dependent on sample size. They were therefore interpreted alongside graphical diagnostics, the magnitude and pattern of departures, and their observed effect on the resulting inference. Table 2 describes the assumption violations in each paper and then puts their severity into context.
4. Subjectivity in inferential reproducibility assessment
An important challenge in the manuscript concerns the inherent subjectivity involved in evaluating assumption violations and selecting alternative models.
The manuscript appropriately incorporates expert judgement and contextual interpretation. However, different statisticians may reasonably disagree regarding:
- whether a violation is “minor” or “substantial”;
- whether a model is sufficiently robust;
- or whether an alternative specification is preferable.
This issue is particularly relevant because the inferential reproducibility classifications partly depend on judgement-based interpretation of:
- diagnostic plots;
- model fit;
- coefficient changes;
- and sensitivity analyses.
The manuscript would benefit from a more explicit discussion of inter-analyst variability and the possibility that different analytical teams could reach different inferential reproducibility conclusions using the same datasets.
The manuscript does acknowledge this issue in the limitations section. However, the point could be more explicitly connected to the interpretation of the binary inferential reproducibility classifications and to the generalisability of the study findings.
We agree that different statisticians may reasonably reach different conclusions when assessing model assumptions and selecting alternative analyses, and this was explicitly acknowledged in the limitations. However, such inter-analyst variability is inherent to statistical modelling and is directly relevant to the concept of inferential reproducibility examined in this study, rather than being introduced solely by our framework. The original discussion considers evidence from multi-analyst studies showing that, even among experienced analysts examining the same data, different defensible analytical choices and conclusions may arise.
We added to the limitations
“Assumption violations were not considered consequential solely because they were present, as their effects depend on their nature, severity, and the context of the analysis. When a violation was identified, expert judgement was used to assess its potential importance, and sensitivity analyses were conducted to determine whether addressing it materially affected the estimates or resulting inferences. This process required judgement in selecting appropriate alternative models and determining whether observed differences were substantively meaningful.”
We do not consider the binary classification of inferential reproducibility a limitation. In statistical practice, statisticians must ultimately determine whether a model provides an adequate basis for the intended inference. The binary inferential reproducibility (Yes/No) classification summarised this overall judgement by indicating whether the original inference remained credible after an assumption violation had been addressed.
The original manuscript provides detailed paper-level reports, referenced in Table 2, describing the diagnostic findings, whether departures were considered minor or substantial, the rationale for each sensitivity analysis, and the basis for the final classification. The original HTML files are also available on GitHub, allowing readers to examine the evidence underlying each decision and form their own judgement.
5. Binary classification of inferential reproducibility may oversimplify uncertainty
The binary Yes/No classification may oversimplify what is inherently a continuum of inferential robustness.
Some classifications appear difficult to interpret because:
- multiple criteria are combined;
- thresholds are partly exploratory;
- and judgement-based interpretation remains involved.
For example:
- 10% changes;
- standardised coefficient differences;
- CI width differences;
- AIC/BIC changes;
- and diagnostic interpretation are all incorporated simultaneously.
The manuscript would benefit from:
- a more explicit hierarchy of decision criteria,
or
- a graded reproducibility framework rather than a binary outcome classification.
Inferential reproducibility is inherently complex, and the framework was developed to reflect the reasoning used in statistical practice. Numerical thresholds were used to support consistent assessment of changes in estimates and precision, but they were interpreted alongside study design, model diagnostics, the nature and severity of any assumption violations, and the suitability of the alternative analysis. The final Yes/No classification therefore represented an overall statistical judgement about whether the original inference remained credible, rather than being determined by any single threshold or a fixed combination of criteria.
6. Justification for inferential thresholds requires further development
The manuscript introduces thresholds such as:
- 10% coefficient changes;
- standardised coefficient differences of 0.1 and 0.2, to define meaningful inferential differences.
While the manuscript acknowledges limitations of percentage-based thresholds and appropriately explores alternative standardised metrics, the theoretical justification for selecting these particular cut-offs remains relatively limited.
Given that the manuscript focuses on methodological rigour and inferential interpretation, stronger justification for these thresholds would improve the conceptual robustness of the framework.
The rationale and limitations of these thresholds were described in the manuscript, including the instability of percentage changes when coefficients are close to zero and the use of standardised differences as an alternative metric. The thresholds were intended as pragmatic decision aids rather than universally validated cut-offs and were interpreted alongside diagnostic findings and expert judgement.
We have added the following clarification to the limitations:
“These pragmatic cut-offs were informed by the research team’s experience in judging whether small differences were substantively meaningful and by conventional benchmarks for small effect sizes. However, further research is needed to validate these thresholds.”
7. Limited contextual knowledge of original studies may influence classifications
The authors appropriately acknowledge that they did not contact original authors and retained original variable selections despite possible overfitting or unclear modelling rationale.
However, this limitation has potentially important implications.
Some reproduced analyses may classify studies as inferentially non-reproducible because:
- important contextual information;
- preprocessing decisions;
- clustering structure;
- transformations;
- or theoretical modelling rationale were unavailable.
This limitation should be discussed more prominently when interpreting inferential reproducibility estimates.
Thank you for this comment. The manuscript already acknowledges that unavailable contextual information may have affected some assessments. However, inferential reproducibility requires sufficient reporting and documentation for an independent analyst to understand whether the model and its assumptions were appropriate. If preprocessing decisions, transformations, clustering structures or modelling rationales cannot be recovered from the article and shared materials, this is directly relevant to inferential reproducibility.
We have added to the limitations
“Accordingly, the individual classifications and overall inferential reproducibility estimate reflect what could be independently assessed from the published articles and shared materials; additional undocumented contextual information may have altered some assessments.”
8. Risk of hindsight optimization in alternative model selection
In several examples, alternative models appear selected after observing diagnostic problems.
This introduces potential concerns regarding:
- researcher degrees of freedom;
- post hoc optimization;
- and analytic subjectivity.
Although the manuscript openly describes these procedures, stronger discussion is needed regarding:
- whether alternative models were pre-specified;
- how competing models were prioritised;
- and whether different analysts might reasonably arrive at different “best” models.
This issue is especially important given the manuscript’s focus on inferential reproducibility.
We do not consider these analyses to constitute post hoc optimisation because they were selected to address specific diagnostic concerns rather than to obtain more favourable results. Although evaluating alternative models after examining the data may increase the risk of overfitting or producing findings that are specific to the observed dataset, this risk was reduced by considering graphical diagnostics, model fit, the suitability of each model for the data, and the consistency of findings across other reasonable analyses.
The exact alternative model cannot be prespecified because the relevant form of misspecification may become evident only after the original model and data have been examined. The study therefore prespecified a structured assessment framework rather than a fixed alternative model for every possible problem. Additional analyses were undertaken only when supported by the diagnostics and were selected to address the specific concern identified.
The aim was not to identify a single “best” model, but to determine whether the original inference remained credible under a reasonable analysis that addressed the identified problem. The rationale for each sensitivity analysis is documented in the paper-level reports (Table 2), and the manuscript acknowledges that another statistician might sometimes select a different, but equally defensible, approach.
Minor Concerns
1. Clarification of terminology
The manuscript occasionally uses:
- reproducibility;
- robustness;
- sensitivity;
- and validity in overlapping ways.
A concise terminology table may improve clarity.
Changes were made to improve clarity. For further detail, see major concern 1.
2. Interpretation of AIC/BIC improvements
Although the manuscript notes that AIC/BIC thresholds should be interpreted in context, some passages could more clearly distinguish improved relative model fit from stronger evidence of substantive or inferential validity. Information criteria primarily reflect relative predictive fit and penalised likelihood rather than direct evidence of model validity. This distinction should be clarified.
We agree that AIC and BIC assess relative support among candidate models and do not independently establish substantive or inferential validity. In the manuscript, they were used alongside graphical assessment, diagnostic findings, theoretical plausibility, model appropriateness, and the resulting estimates. We have nevertheless added a paragraph to the Discussion, following the discussion of influential observations, to clarify why information criteria should not be interpreted in isolation.
“Papers 32 and 46 also illustrate why models should not be selected solely on the basis of AIC or BIC (Table 2). Although these criteria compare the relative fit of candidate models while penalising model complexity, the model with the lowest value is not necessarily the most scientifically plausible. In these examples, polynomial terms initially appeared to improve model fit, but graphical assessment showed that the apparent relationships were strongly influenced by extreme observations. A higher-order polynomial may similarly improve AIC or BIC by closely following features specific to the observed dataset while representing a relationship that is difficult to justify or unlikely to generalise. Information criteria should therefore be interpreted alongside graphical assessment, residual diagnostics, theoretical plausibility, and the influence of individual observations.”
3. Reporting burden and readability
The manuscript is highly detailed and methodologically dense. Although this improves transparency, some sections could be shortened or reorganized for readability, particularly the extended methodological descriptions.
We appreciate the reviewer’s concern regarding readability. However, we do not agree that the methodological descriptions should be substantially shortened. Assessing inferential reproducibility requires several analytical decisions that are not yet governed by a single established framework, including assessing model assumptions, selecting targeted sensitivity analyses, and determining whether results from alternative analyses can be validly compared. Sufficient methodological detail is therefore necessary to make the statistical reasoning transparent and reproducible. To support readability, we have organised the methodological material using clear headings and subheadings, allowing readers to navigate directly to the sections most relevant to them.
4. Discussion could better distinguish educational versus structural problems
The conclusion appropriately emphasises statistical education. However, inferential reproducibility problems may also reflect:
- publication incentives;
- inadequate peer review resources;
- insufficient methodological collaboration;
- and broader systemic pressures within academic publishing.
This broader context could be acknowledged more explicitly.
Thank you for your comment; we have added a paragraph discussing structural problems in the conclusion.
“However, inferential reproducibility problems should not be attributed solely to gaps in individual researchers’ statistical knowledge. Publication incentives, including pressure to publish, and broader systemic pressures may limit the time, resources, and support available for careful model development, diagnostic assessment, and transparent reporting [73]. Limited peer-review capacity and inconsistent access to methodological expertise may further reduce opportunities for statistical problems to be identified before publication [74, 75].”
Suggestions for Improvement
1. Clarify the conceptual definition of inferential reproducibility and distinguish it more explicitly from robustness and sensitivity analysis.
Changes were made to improve clarity. For further detail, see major concern 1.
2. Emphasize more clearly that the inferential reproducibility estimates are exploratory given the relatively small sample size.
The small sample size was reiterated in the limitations section. For details, see major concern 2.
3. Provide a stronger theoretical justification for the selected inferential thresholds (10%, 0.1, 0.2).
We have added clarification to the limitations. For further detail, see major concern 6.
4. Expand discussion regarding robustness of linear regression to moderate assumption violations.
We have added paragraphs discussing this in both the discussion and limitations; for more detail, see major concern 1 and 4.
5. Discuss more explicitly the role of subjective judgement and potential inter-analyst variability in inferential reproducibility assessments.
We added a discussion point to the limitations. For further detail, see major concern 4.
6. Clarify how alternative models were selected and whether multiple reasonable modelling approaches could coexist.
We added further detail to clarify how alternative models were selected. For further detail, see major concern 1.
7. Expand discussion regarding the limitations created by unavailable contextual information from original authors.
We have added a point to the limitations section. For further detail, see major concern 7.
8. Consider including a summary decision flowchart explaining how inferential reproducibility classifications were determined.
We have included a flowchart (Figure 2) to clarify understanding of the inferential reproducibility process.
Overall Evaluation
Thank you for your review. We have clarified the methodological framework and the distinction between inferential reproducibility, robustness, and sensitivity analysis. We have also expanded the explanation of how assumption violations, numerical thresholds, information criteria, and expert judgement informed the final classifications. Additional revisions address the interpretation of the binary classification, model-selection subjectivity, inter-analyst variability, and the study’s limitations. Detailed responses to each point are provided above.
Reviewer 2: Konrad Hinsen
1. Page 6 describes the HTML reports produced by the authors for each of the examined papers, but does not provide a link to them. They are in fact accessible from Ref. 19, which also contains all the code for this work. Ref. 19. is referred to in several places in the paper for different reasons. I think it would be more helpful to the reader to discuss all the material available in Ref. 19 in a single place, given that it is the central repository for the code and data backing up this study, rather than a distinct work that looks like just one cited reference out of many others.
While not directly in the manuscript, the data availability statement is available on medRxiv and has not been duplicated, as this metadata is usually included on the paper’s first page when it is published, along with conflict of interest, etc.
Data Availability
The data and a reproducible R Quarto file used to produce this paper, including tables, figures, and code, have been stored in a GitHub repository and can be cited using the Zenodo DOI.
https://github.com/Lee-V-Jones/Reproducibility
https://doi.org/10.5281/zenodo.19448969
However, for ease of access, we have added a direct link in the text to the 95 HTML reports. “The reports are publicly available at https://github.com/Lee-V-Jones/Reproducibility/blob/main/README.md”
2. Acronyms are mostly defined at first use, but I found a few exceptions: AIC, BIC, DFFITS, DFBETAS, COVRATIO. Also, GAM appears in the caption of Figure 1, on page 4, before being defined in the main text on page 5.
Thanks for picking this up. This was an oversight: we defined acronyms for GAM, AIC, BIC, ANOVA, DFFITS, DFBETAS, and COVRATIO the first time they were used. We have edited Figure 1 to define Generalised Additive Models (GAM).





