Evaluated works

Research Articles

  • Article

    Avihay Cohen, Blanka Ivanović, Anastasiia Iarkaeva, Vladislav Nachev, Evgeny Bobrov

    Data sharing is increasingly expected by funders, journals, institutions, and other actors in the research system. This expectation is based primarily on the assumption that some of the shared data will be reused. Reuse can take many forms and have many purposes, but currently the only way to detect reuse of data at scale is to screen the research literature. We used the Data Citation Corpus to detect instances in which datasets shared by researchers from our biomedical research institution had been referenced, which we interpreted as indicative of reuse. We observed that 4.9% of datasets shared between 2020 and 2023 had been referenced by February 2025. Overall, 175 Charité datasets from this period were reused, which had been referenced 1497 times. The large majority of reused datasets were from ‘omics’ fields, which generate particularly structured and standardized data. Data from humans had a higher probability of reuse, as well as COVID-19-related data. Dataset properties indicative of technical reusability as DOIs and CC licenses had no positive influence on reuse probability. We conducted a complimentary analysis of datasets shared alongside data articles. Here, we observed that a major fraction of references to data articles were in fact referring to shared datasets. Based on a sample, we extrapolated that indirect data citations account for an additional approximately 846 reuse cases. Unlike data references, indirect data citations often referred to datasets in disciplinary repositories from fields outside of ‘omics’, as well as general-purpose repositories. Our study shows that data are reused on a large scale, but at the same time reuse of data in many fields is limited, especially if data are deposited in general-purpose repositories. Data is reused most if it is shared with disciplinary metadata and/or described by data articles. Our methods do not allow us to map data reuse in its entirety, and further large-scale studies on the determinants, temporal patterns and purposes of data reuse are needed.

    July 22, 2026

  • Article

    Andrea Cadeddu, Alessandro Chessa, Vincenzo De Leo, Gianni Fenu, Francesco Osborne, Diego Reforgiato Recupero, Angelo Salatino, Luca Secchi

    The United Nations' Sustainable Development Goals (SDGs) provide a globally recognised framework for addressing critical societal, environmental, and economic challenges. Recent developments in natural language processing (NLP) and large language models (LLMs) have facilitated the automatic classification of textual data according to their relevance to specific SDGs. Nevertheless, in many applications, it is equally important to determine the directionality of this relevance; that is, to assess whether the described impact is positive, neutral, or negative. To tackle this challenge, we propose the novel task of SDG polarity detection, which assesses whether a text segment indicates progress toward a specific SDG or conveys an intention to achieve such progress. To support research in this area, we introduce SDG-POD, a benchmark dataset designed specifically for this task, combining original and synthetically generated data. We perform a comprehensive evaluation using six state-of-the-art large LLMs, considering both zero-shot and fine-tuned configurations. Our results suggest that the task remains challenging for the current generation of LLMs. Nevertheless, some fine-tuned models, particularly QWQ-32B, achieve good performance, especially on specific Sustainable Development Goals such as SDG-9 (Industry, Innovation and Infrastructure), SDG-12 (Responsible Consumption and Production), and SDG-15 (Life on Land). Furthermore, we demonstrate that augmenting the fine-tuning dataset with synthetically generated examples yields improved model performance on this task. This result highlights the effectiveness of data enrichment techniques in addressing the challenges of this resource-constrained domain. This work advances the methodological toolkit for sustainability monitoring and provides actionable insights into the development of efficient, high-performing polarity detection systems.

    July 10, 2026

  • Article

    Adrian Barnett, Matt Spick

    PLOS journals allow authors of accepted articles to choose whether the peer review will be openly published alongside the article. This creates an observational study to examine the characteristics of authors and articles that more often choose open peer review, and whether open review is associated with measures of article quality. We examined over 115,000 PLOS articles and estimated what characteristics of the articles were associated with open peer review. We also examined if open peer review was associated with the subsequent retraction of the article and the number of citations.

    Forty percent of articles chose open peer review. Authors from the UK, France, the Netherlands, and Ethiopia were more likely to choose open peer review. In contrast, authors from Saudi Arabia, the Republic of Korea, Pakistan, Poland, and China were less likely to choose open review. Authors with an edu email were less likely to choose open review, whilst authors with a gmail were more likely. Articles with open peer review were less likely to be retracted (adjusted hazard ratio = 0.75, 95% CI 0.60 to 0.93) and had more citations on average (adjusted rate ratio = 1.07, 95% CI 1.05 to 1.09).

    In this exploratory study, we found clear differences in participation in open reviews with strong differences between countries. Authors who have confidence in their article and who engage in other open science practices may be more likely to chose open peer review, making this choice an indicator of article quality.

    July 1, 2026

  • Article

    Lutz Bornmann, Christian Leibel

    Citation analysis is widely used in research evaluation to assess the impact of scientific papers. These analyses rest on the assumption that citation decisions by authors are accurate, representing the flow of knowledge from cited to citing papers. However, in practice, researchers often cite for reasons that are not related to the fact that there has been (intellectual) input from previous papers. Citations made for rhetorical reasons or without reading the cited work compromise the value of citations as instrument for research evaluation. Past research on threats to the accuracy of citations has mainly focused on citation bias as the primary concern. In this paper, we argue that citation noise - the undesirable variance in citation decisions - represents an equally critical but underexplored challenge in citation analysis. We define and differentiate two types of citation noise: citation level noise and citation pattern noise. Each type of noise is described in terms of how it arises and the specific ways it can undermine the validity of citation-based research assessments. By conceptually differing citation noise from citation accuracy and citation bias, we propose a framework for the foundation of citation analysis. We discuss strategies and interventions to minimize citation noise, aiming to improve the reliability and validity of citation analysis in research evaluation. We recommend that the current professional reform movement in research evaluation such as the Coalition for Advancing Research Assessment (CoARA) pick up these strategies and interventions as an additional building block for careful, responsible use of bibliometric indicators in research evaluation.

    June 26, 2026

  • Article

    Minhyuk Park, Haotian Yi, Tandy Warnow, George Chacko

    The global science literature, represented as a network with articles as nodes and citations as edges, is a rich artifact for scientometric studies. The structure of this network depends on its count of nodes and distribution of edges. We are interested in how the network has evolved from its origin to its present structure. Since extant theories of citation do not offer much in the way of quantitative explanation, we use a modeling approach that generates synthetic networks. Specifically, we have developed an idealized agent-based model of citations (SASCA-ReS) that can generate synthetic networks of size over 200 million nodes, which is comparable with the size of today’s science literature. This model allows us to reason in an artificial world and to identify patterns of citation that may explain real-world scenarios. We report results from simulations under this model with different parameter settings and at various scales to explore counterfactual and hypothetical scenarios.

    June 16, 2026

  • Article

    Juan Pablo Bascur, Rodrigo Costas, Suzan Verberne

    Traditional science maps visualize topics by clustering documents within a network, but they are inherently biased toward clustering certain topics over others. If these topics could be chosen, then the science maps could be tailored for different needs. In this paper, we explore the extent to which the topic bias of a science map can be changed by choosing different data sources to build the document network. We analyze this by evaluating the clustering effectiveness of several topic categories over two sources that are traditionally used for the creation of science maps (citations and text similarity) and six non-traditional data sources, which we found favor different kinds of topics: Health issues for Facebook users, biotechnology topics for patent families, government and social issues for policy documents, food topics for Twitter conversations, nursing topics for Twitter users, and geographical entities for document authors (the favoring in this latter source was particularly strong). Our results show that diverse data sources can be used to control topic bias, which opens up the possibility of creating science maps tailored for different needs.

    June 11, 2026

  • Article

    Hugh Shanahan, Louise Bezuidenhout

    Preprint services now play a key role in disseminating research across a wide range of domains. In this paper we examine where a set of 64 preprint services, collated by ASAPbio, are being physically hosted and the type of Internet Service Provider (ISP) they are using. In addition to this access to these services from 106 territories was simulated using Virtual Private Networks (VPNs). We find that the majority of services (47/64) are physically based in the USA, despite instances where the Top Level Domain (TLD) indicates a different country. In addition to this 56/64 of the services are being hosted by commercial ISPs with 43/64 being hosted by Cloudflare, Google LLC or Amazon. Poorer countries are more likely to encounter an error when attempting a URL corresponding to the landing page of one of the preprint services collated by ASAPbio than more wealthy countries. Sites that are physically located in Low or Middle Income Countries have similar accessibilities and may be better. We draw some overall conclusions about the potential frailty of these services with respect to the technical decisions made here.

    June 5, 2026

  • Article

    Wolfgang Kaltenbrunner, Andrea Chiarelli, Andrea Reyes Elizondo, Stephen Pinfield, Ludo Waltman, André Brasil

    We analyse focus group discussions and free-text survey responses from a multi-stakeholder consultation conducted after the October 2023 publication of the proposal Towards Responsible Publishing by cOAlition S. The proposal calls for a systemic reform of scholarly communication by reducing barriers to knowledge dissemination, promoting early sharing of outputs through preprints, and shifting peer review to an open, post-publication model. We focus on how different stakeholder groups –such as researchers, infrastructure providers, academic institutions, and publishers –perceive obstacles to the large-scale, coordinated reform envisioned in the proposal. We interpret these accounts as articulations of collective action problems, shaped by entrenchment of many actors in existing academic reward systems and established commercial revenue models that make transitions toward a more economically sustainable scholarly communication system difficult, even where many actors see the principal need for change. This approach highlights the extent to which stakeholder perspectives align or conflict. It also underscores the performative nature of discourse about collective action problems in scholarly communication: by articulating challenges to reform, participants simultaneously construct, reinforce, or contest their own roles within the system, which directly influences their collective capacity to act.

    June 3, 2026

  • Research plan

    Savannah C. Lewis, Alexa M. Tullett

    Scientific integrity depends on ethical and transparent research practices, yet surveys reveal that many researchers engage in questionable research practices (QRPs), ranging from minor issues (e.g., unclear preregistration) to major misconduct (e.g., data fabrication). This study investigates how people perceive researchers who commit QRPs (investigators) compared to those who report them (inspectors). Participants (N ≈ 566) will read three hypothetical vignettes describing an investigator engaging in a QRP and an inspector who reports it. Participants will then evaluate both researchers on trustworthiness and likeability. They will also rate the perceived trueness of the original finding, and compare both researchers on a range of positive attributes. Thus, the overall design is a 2 (role: investigator vs. inspector) x 3 (QRP severity: minor vs. moderate vs. major) fully within-subjects design. Multilevel models will test whether perceptions vary by role and QRP severity. This research will deepen our understanding of how accountability in science is socially evaluated, and how the severity of misconduct shapes views of both those who commit and those who call out QRPs.

    May 28, 2026

  • Article

    Alex Hulkes

    All other things being equal, two-stage research funding processes that involve the initial submission of a relatively short proposal containing only minimal information (typically referred to as an outline proposal) followed by a preparation and submission of a full proposal will on average require less effort from their applicants than will single-stage processes. But it is likely that a funder operating a process which begins with a lower-effort outline stage will receive more applications than they might have expected to see had applicants been required initially to prepare a full proposal in a single-stage process. The net effect of these interacting and competing influences on the overall effort required in, and efficiency of, a funding process is not currently known and so is investigated in this work using an Agent-Based Modelling approach. The results of this model suggest that while the number of applications submitted will indeed increase, perhaps by as much as 40%, if two-stage processes are used, the level of applicant effort per unit output (that is, unit of funding awarded or number of awards made) may be reduced by around 15% to 20%. A weaker but more general interpretation, that does not rely so much on the specifics of the model, is that substantial increases in demand arising from use of outline processes might still come with an overall decrease in applicant effort. A reasonable conclusion is that more extensive use of two-stage research funding processes may lead to significant cost savings.

    May 12, 2026