Evaluated works
Research Articles
Article
Dorothea Strecker, Heinz Pampel, Jonas Höfting
This article presents the results of a survey conducted in 2024 among research performing organizations (RPOs) in Germany on how they collect data on publication costs. Of the 583 invitees, 258 (44.3%) completed the questionnaire. This survey is the first comprehensive study on the recording of publication costs at RPOs in Germany. The results show that the majority of surveyed RPOs recorded publication costs at least in part. However, procedures in this regard were often non-binding. Respondents' ratings of the reliability of the collection of data on publication costs varied by the source of publication funding. Eighty percent of respondents rated the contribution of collecting data on publication costs to shaping the open access transformation as "very important" or "important." Yet, these data were used as a basis for strategic decisions in only 59% of the surveyed RPOs. Moreover, most respondents considered the implementation of an information budget at their institutions by 2025 unlikely. We discuss the implications of these findings for the open access transformation.
August 13, 2026
Article
Daniel W. Hook
The debate about scholarly knowledge infrastructure has long been framed as a contest between openness and commercial enclosure. This framing distorts both policy and practice. The real tension lies between the persistent cost of producing and refining structured metadata under deep technological friction, and the differentiated demands distinct communities place on data quality, focus and granularity. We introduce the innovation annulus: the zone between freely available structured data and the advancing frontier of commercially refined knowledge products. This zone is a permanent, functional feature of the ecosystem -- not a pathology to eliminate. By analogy with the efficient market hypothesis, its width measures production inefficiency, set by the interplay of friction and demand. Artificial intelligence reshapes the annulus, lowering barriers to basic structuring, raising the threshold at which refinement adds value, and introducing systemic risks through unprovenanced AI-derived metadata. CRediT contributions, funding acknowledgements and AI disclosure statements illustrate the annulus lifecycle. Governance should calibrate the annulus, not abolish it: thin enough to serve research efficiently, wide enough to sustain innovation. A formal welfare framework, analogous to the Nordhaus optimal patent life, characterises the trade-offs and yields testable predictions. The Barcelona Declaration offers a promising forum for boundary governance.
August 12, 2026
Article
Lee Jones, Adrian Barnett, Gunter Hartel, Dimitrios Vagenas
Background: In health research, variability in modelling decisions can lead to different conclusions even when the same data are analysed, a challenge known as inferential reproducibility. In linear regression analyses, incorrect handling of key assumptions, such as normality of the residuals and linearity, can undermine reproducibility. This study examines how violations of these assumptions influence inferential conclusions when the same data are reanalysed.
Methods: We randomly sampled 95 health-related PLOS ONE papers from 2019 that reported linear regression in their methods. Data were available for 43 papers, and 20 were assessed for computational reproducibility, with three models per paper evaluated. The 14 papers that included a model at least partially computationally reproduced were then examined for inferential reproducibility. To assess the impact of assumption violations, differences in coefficients, 95% confidence intervals, and model fit were compared.
Results: Of the fourteen papers assessed, only three were inferentially reproducible. The most frequently violated assumptions were normality and independence, each occurring in eight papers. Violations of independence were particularly consequential and were commonly associated with inferential failure. Although reproduced analyses often retained the same binary statistical significance classification as the original studies, confidence intervals were frequently wider, indicating greater uncertainty and reduced precision. Such uncertainty may affect the interpretation of results and, in turn, influence treatment decisions and clinical practice.
Conclusion: Our findings demonstrate that substantial violations of key modelling assumptions often went undetected by authors and peer reviewers and, in many cases, were associated with inferential reproducibility failure. This highlights the need for stronger statistical education and greater transparency in modelling decisions. Rather than applying rigid or misinformed rules, such as incorrectly testing the normality of the outcome variable, researchers should adopt modelling frameworks guided by the research question and the study design. When assumptions are violated, appropriate alternatives, such as robust methods, bootstrapping, generalized linear models, or mixed-effects models, should be considered. Given that assumption violations were common even in relatively simple regression models, early and sustained collaboration with statisticians is critical for supporting robust, defensible, and clinically meaningful conclusions.
August 10, 2026
Article
Simon Porter, Daniel Hook
Bibliographic data is a rich source of information that goes beyond the use cases of location and citation -- it also encodes both cultural and technological context. For most of its existence, the scholarly record has changed slowly and hence provides an opportunity to gain insight through its reflection of the cultural norms of the research community over the last four centuries. While it is often difficult to distinguish the originating driver of change, it is still valuable to consider the motivating influences that have led to changes in the structure of the scholarly record. An "initial era" is identified during which initials were used in preference to full names by authors on scholarly communications. Causes of the emergence and demise of this era are considered as well as the implications of this era on research culture and practice.
August 5, 2026
Article
Zhicheng Lin
Scientific communication faces a dual crisis: exponential publication growth—now accelerated by AI-assisted writing—overwhelms human readers and reviewers, while fragmented research practices block automated synthesis. The behavioral and social sciences in particular suffer from incomparable stimulus databases, jingle–jangle measurement fallacies (same label, distinct constructs; different labels, same construct), and contextual blindness that conceals effect heterogeneity. Current AI tools can summarize papers but cannot synthesize findings across incommensurable studies; they also risk amplifying biases when trained on unstructured, unverified text. I propose restructuring scientific papers for dual audiences: front-loaded narratives for time-pressed human readers, paired with research-object packages containing executable code, semantic annotations, and tidy trial-level data. This design makes papers queryable research environments: readers can interrogate data and probe analytic choices in real time, while research-object packages enable automated verification and AI-assisted peer review grounded in executable evidence rather than narrative claims. Such papers become nodes in continuously updated evidence networks: each publication automatically contributes effect sizes to living, versioned meta-analyses, with corrections and retractions propagating through dependent analyses. Widespread adoption will require institutional recognition of structured documentation as essential scholarly output and computational infrastructure that serves both human comprehension and machine analysis.
July 28, 2026
Article
Soohong Eum
Comprehensive and reliable funding information is essential for evaluating public research investment, yet funding metadata remain fragmented across heterogeneous data sources maintained by funders and bibliographic platforms. This study examines the interoperability between a national funder database, the National Science and Technology Information Service (NTIS) of South Korea, and two bibliographic sources, Web of Science (WoS) and OpenAlex, in order to assess the quality and coverage of funded publication data. Using a multi-step metadata matching procedure, more than 99% of NTIS records were successfully linked to corresponding documents in WoS and OpenAlex, demonstrating that bibliographic metadata in NTIS records are generally complete and accurate despite incomplete coverage of persistent identifiers such as DOIs. However, limitations are identified in funding-specific metadata, particularly in the centrally assigned contribution rate, which does not account for non-Korean or non-governmental funding and may bias national statistics. Comparison across data sources reveals substantial overlap but also notable differences in coverage and granularity. These differences reflect the distinct collection mechanisms of funders’ databases and bibliographic sources and give rise to complementarities that can be exploited to obtain a more comprehensive picture of funded research outputs and their associated funding sources. The findings carry broader implications for funding agencies seeking to improve interoperability, metadata standardisation, and the analytical utility of research information systems.
July 23, 2026
Article
Avihay Cohen, Blanka Ivanović, Anastasiia Iarkaeva, Vladislav Nachev, Evgeny Bobrov
Data sharing is increasingly expected by funders, journals, institutions, and other actors in the research system. This expectation is based primarily on the assumption that some of the shared data will be reused. Reuse can take many forms and have many purposes, but currently the only way to detect reuse of data at scale is to screen the research literature. We used the Data Citation Corpus to detect instances in which datasets shared by researchers from our biomedical research institution had been referenced, which we interpreted as indicative of reuse. We observed that 4.9% of datasets shared between 2020 and 2023 had been referenced by February 2025. Overall, 175 Charité datasets from this period were reused, which had been referenced 1497 times. The large majority of reused datasets were from ‘omics’ fields, which generate particularly structured and standardized data. Data from humans had a higher probability of reuse, as well as COVID-19-related data. Dataset properties indicative of technical reusability as DOIs and CC licenses had no positive influence on reuse probability. We conducted a complimentary analysis of datasets shared alongside data articles. Here, we observed that a major fraction of references to data articles were in fact referring to shared datasets. Based on a sample, we extrapolated that indirect data citations account for an additional approximately 846 reuse cases. Unlike data references, indirect data citations often referred to datasets in disciplinary repositories from fields outside of ‘omics’, as well as general-purpose repositories. Our study shows that data are reused on a large scale, but at the same time reuse of data in many fields is limited, especially if data are deposited in general-purpose repositories. Data is reused most if it is shared with disciplinary metadata and/or described by data articles. Our methods do not allow us to map data reuse in its entirety, and further large-scale studies on the determinants, temporal patterns and purposes of data reuse are needed.
July 22, 2026
Article
Andrea Cadeddu, Alessandro Chessa, Vincenzo De Leo, Gianni Fenu, Francesco Osborne, Diego Reforgiato Recupero, Angelo Salatino, Luca Secchi
The United Nations' Sustainable Development Goals (SDGs) provide a globally recognised framework for addressing critical societal, environmental, and economic challenges. Recent developments in natural language processing (NLP) and large language models (LLMs) have facilitated the automatic classification of textual data according to their relevance to specific SDGs. Nevertheless, in many applications, it is equally important to determine the directionality of this relevance; that is, to assess whether the described impact is positive, neutral, or negative. To tackle this challenge, we propose the novel task of SDG polarity detection, which assesses whether a text segment indicates progress toward a specific SDG or conveys an intention to achieve such progress. To support research in this area, we introduce SDG-POD, a benchmark dataset designed specifically for this task, combining original and synthetically generated data. We perform a comprehensive evaluation using six state-of-the-art large LLMs, considering both zero-shot and fine-tuned configurations. Our results suggest that the task remains challenging for the current generation of LLMs. Nevertheless, some fine-tuned models, particularly QWQ-32B, achieve good performance, especially on specific Sustainable Development Goals such as SDG-9 (Industry, Innovation and Infrastructure), SDG-12 (Responsible Consumption and Production), and SDG-15 (Life on Land). Furthermore, we demonstrate that augmenting the fine-tuning dataset with synthetically generated examples yields improved model performance on this task. This result highlights the effectiveness of data enrichment techniques in addressing the challenges of this resource-constrained domain. This work advances the methodological toolkit for sustainability monitoring and provides actionable insights into the development of efficient, high-performing polarity detection systems.
July 10, 2026
Article
Adrian Barnett, Matt Spick
PLOS journals allow authors of accepted articles to choose whether the peer review will be openly published alongside the article. This creates an observational study to examine the characteristics of authors and articles that more often choose open peer review, and whether open review is associated with measures of article quality. We examined over 115,000 PLOS articles and estimated what characteristics of the articles were associated with open peer review. We also examined if open peer review was associated with the subsequent retraction of the article and the number of citations.
Forty percent of articles chose open peer review. Authors from the UK, France, the Netherlands, and Ethiopia were more likely to choose open peer review. In contrast, authors from Saudi Arabia, the Republic of Korea, Pakistan, Poland, and China were less likely to choose open review. Authors with an edu email were less likely to choose open review, whilst authors with a gmail were more likely. Articles with open peer review were less likely to be retracted (adjusted hazard ratio = 0.75, 95% CI 0.60 to 0.93) and had more citations on average (adjusted rate ratio = 1.07, 95% CI 1.05 to 1.09).
In this exploratory study, we found clear differences in participation in open reviews with strong differences between countries. Authors who have confidence in their article and who engage in other open science practices may be more likely to chose open peer review, making this choice an indicator of article quality.
July 1, 2026
Article
Lutz Bornmann, Christian Leibel
Citation analysis is widely used in research evaluation to assess the impact of scientific papers. These analyses rest on the assumption that citation decisions by authors are accurate, representing the flow of knowledge from cited to citing papers. However, in practice, researchers often cite for reasons that are not related to the fact that there has been (intellectual) input from previous papers. Citations made for rhetorical reasons or without reading the cited work compromise the value of citations as instrument for research evaluation. Past research on threats to the accuracy of citations has mainly focused on citation bias as the primary concern. In this paper, we argue that citation noise - the undesirable variance in citation decisions - represents an equally critical but underexplored challenge in citation analysis. We define and differentiate two types of citation noise: citation level noise and citation pattern noise. Each type of noise is described in terms of how it arises and the specific ways it can undermine the validity of citation-based research assessments. By conceptually differing citation noise from citation accuracy and citation bias, we propose a framework for the foundation of citation analysis. We discuss strategies and interventions to minimize citation noise, aiming to improve the reliability and validity of citation analysis in research evaluation. We recommend that the current professional reform movement in research evaluation such as the Coalition for Advancing Research Assessment (CoARA) pick up these strategies and interventions as an additional building block for careful, responsible use of bibliometric indicators in research evaluation.
June 26, 2026











