Published at MetaROR
July 28, 2026
Table of contents
Restructuring scientific papers for human and AI readers
1 Department of Psychology, Yonsei University, Seoul 03722, Republic of Korea
Originally published on May 2, 2026 at:
Abstract
Scientific communication faces a dual crisis: exponential publication growth—now accelerated by AI-assisted writing—overwhelms human readers and reviewers, while fragmented research practices block automated synthesis. The behavioral and social sciences in particular suffer from incomparable stimulus databases, jingle–jangle measurement fallacies (same label, distinct constructs; different labels, same construct), and contextual blindness that conceals effect heterogeneity. Current AI tools can summarize papers but cannot synthesize findings across incommensurable studies; they also risk amplifying biases when trained on unstructured, unverified text. I propose restructuring scientific papers for dual audiences: front-loaded narratives for time-pressed human readers, paired with research-object packages containing executable code, semantic annotations, and tidy trial-level data. This design makes papers queryable research environments: readers can interrogate data and probe analytic choices in real time, while research-object packages enable automated verification and AI-assisted peer review grounded in executable evidence rather than narrative claims. Such papers become nodes in continuously updated evidence networks: each publication automatically contributes effect sizes to living, versioned meta-analyses, with corrections and retractions propagating through dependent analyses. Widespread adoption will require institutional recognition of structured documentation as essential scholarly output and computational infrastructure that serves both human comprehension and machine analysis.
Introduction
Publication output doubles roughly every 17 years1, reaching 3.3 million articles in 20222, while researchers spend ever less time on each paper3. AI tools aid summarization and question answering4 but cannot solve the deeper challenge of knowledge integration when the underlying literature is incoherent and inaccessible. In the behavioral and social sciences, findings are fragmented by incomparable measures5, bespoke materials and stimuli6,7, and poorly documented participant demographics8. Stimuli, code, and data9 are often unavailable; when shared, they are frequently incomplete, undocumented, or too coarsely aggregated for verification and reuse10. Knowledge cannot accumulate from such fragments.
To address this dual crisis of volume and fragmentation, and in response to a recent U.S. National Academies report calling for ontology infrastructure to coordinate the behavioral sciences11, I argue that the scientific paper must be rebuilt for two audiences: human readers and AI systems. For humans, this requires front-loading key findings in accessible prose for time-pressed researchers and turning papers into layered knowledge bases where readers can interrogate data, rerun analyses, and access technical details without wading through dense text. For AI systems, the main text, data, code, and protocols must be structured in machine-readable formats that enable automated analysis, comparison, and synthesis.
Here “AI system” denotes an LLM‑based pipeline that combines generative models with tools, external memory, and deterministic analysis modules, rather than a standalone language model. I use “Generative AI” for the broader class of models that produce text, code, or images; “AI agent” for an AI system delegated to perform multi-step tool use within scoped task limits (e.g., searching papers, executing code, generating audit reports); and “AI tools” for user-facing utilities such as summarizers or citation matchers that do not involve delegated task execution. Currently, such pipelines convert source files into segmented text, index those segments with vector and keyword retrieval, and call tools for tasks such as code execution under LLM-orchestrated planning12. Each step introduces potential errors: PDF parsing loses reading order, equations, and metadata; table extraction remains brittle; figure-panel parsing and citation disambiguation likewise rely on inference from rendered pages rather than typed structure13. The structured package proposed here focuses on whether retrieval and tool calls operate over verified research objects or over lossy reconstructions of prose. Some publishers already provide JATS XML or HTML serializations alongside the rendered PDF, which reduces parsing errors at the article-text level; the more consequential gap is that the research objects on which verification depends—data, code, stimuli, protocols, schemas, and provenance—are often absent or weakly linked.
This Perspective extends two decades of work on machine-actionable scholarship. Data, metadata, tools, and workflows must be findable, accessible, interoperable, and reusable (FAIR)14,15. The FORCE11 community and the “Beyond the PDF” movement argued that articles should treat data, software, and protocols as integral research objects rather than appendages; the Research Object (RO) architecture formalized how to bundle them with provenance and attribution16,17. Implementation efforts toward integrated, executable publications include RO-Crate and Frictionless Data (which provide standardized packaging), GigaScience/GigaDB, eLife’s executable articles, EBRAINS Live Papers, and RIO nanopublications15,18-23. At the governance and infrastructure level, the Research Data Alliance, GO FAIR, and the Barcelona Declaration have advanced complementary efforts to coordinate data-sharing infrastructure, converge on FAIR implementation, and make open research information the default24-26. Psychology has not been a bystander: preregistration, Registered Reports, and reproducible code pipelines aim to reform how behavioral research is conducted and reported10,27-30.
Packaging standards provide a substrate for meaning but do not, by themselves, supply the domain-specific semantic layer. RO-Crate can bundle a stimulus set; it does not specify which persistent identifier scheme applies to affective images, which normative dimensions must be recorded, or whether the construct labeled “executive function” in one container refers to the same phenomenon as in another15,31-33. Frictionless Data can describe a table’s schema; it cannot detect that two studies operationalize “resilience” with instruments that share a name but measure different things5,18,34. The result is that even with access to the best of these platforms, a researcher still cannot query across studies—asking which findings used positively valenced stimuli with arousal ratings above 6, or which “resilience” measures share empirical variance with neuroticism rather than with each other. This is not because data are unavailable, but because the missing specification layer—stimulus identifiers, construct ontologies, trial-level observations with demographic and contextual moderators—has not yet been defined for behavioral science35-39. Openness reforms addressed process transparency; they did not address this taxonomic problem, because it cannot be resolved by sharing data alone10,27. The aim here is not to introduce another platform or standard but to specify how existing tools—containers, ontologies, data standards, and registries—can fill that missing layer for behavioral and psychological science, defining what AI systems need to verify, query, compare, and synthesize findings rather than merely retrieve and paraphrase them.
Against that background, this Perspective makes four contributions. First, it proposes a dual‑audience paper architecture: a front‑loaded narrative stating evidence strength, scope, and paths to verification, paired with a research-object package whose code, data, materials, and metadata are central parts of the contribution. Second, it specifies a minimum viable adoption profile for that package: the elements that existing tools and standards—containers, standardized stimuli, persistent identifiers, ontological mappings of constructs, and tidy trial‑level data with inclusive demographic and contextual coding—must capture for behavioral science specifically. Third, it introduces a publication‑layer jingle–jangle audit in which semantic and statistical checks on construct labels become routine quality control; this requires not just structured data but a theory of construct identity. Fourth, it specifies what individual papers must export—effect sizes, construct mappings, stimulus metadata, sample and context moderators, and quality indicators—to function as nodes in continuously updated, quality-stratified living evidence networks.
The architecture’s full executable form suits empirical, data-rich fields most directly; in interpretive disciplines the relevant research objects shift to corpora, editions, annotation schemes, and the provenance of coding or editorial decisions, but the underlying principle— claims linked to their inspectable evidence, methods, and revision history—holds across fields, implemented through whatever objects each discipline treats as its evidence.
The volume–fragmentation spiral
The volume–fragmentation spiral has produced a cascade of institutional failures, beginning with quality control. Exponential growth in publications and preprints has overwhelmed traditional peer review40, which no longer provides a reliable signal of quality and relevance. Researchers skim abstracts, abandon papers mid-reading, and retreat to secondary summaries. Meanwhile, critical data are buried in unstructured prose, trapped behind paywalls, and encoded in formats that resist computation. This friction may incidentally slow the automated propagation of errors and the inappropriate reuse of sensitive data, but it is a crude safeguard; structured, machine-readable archives could instead pair easier access with explicit governance and quality checks, enabling more reliable synthesis.
These institutional failures both reflect and deepen epistemic fragmentation. Experimental psychologists, for example, routinely deploy proprietary or poorly documented stimulus sets—images, videos, vignettes, auditory clips—that differ in valence, arousal, cultural reference, and perceptual salience35,36. Even well-validated sets are seldom cross-validated against one another; examples include the many affective picture databases: the International Affective Picture System (IAPS)41, Open Affective Standardized Image Set (OASIS)42, Geneva Affective Picture Database (GAPED)43, Nencki Affective Pictures System (NAPS)44, Complex Affective Scene Set (COMPASS)45, International Affective Virtual Reality System (IAVRS)46, and dozens of culturally specific successors47. Effects may then be driven by stimulus-specific confounds rather than the intended construct. Without standardized, cross-validated stimulus sets, studies using different databases or materials become incomparable, forcing meta-analysts to hand-code or exclude findings—a laborious, error-prone process48.
The measurement landscape is equally fragmented. Psychology and related behavioral and social sciences are rife with “jingle–jangle” fallacies: identical names for fundamentally different phenomena (the jingle fallacy49) and different names for the same construct (the jangle fallacy50). “Flourishing,” for instance, bundles conflicting theoretical approaches while preserving nominal unity51. “Executive function” spans working memory, cognitive flexibility, and inhibitory control32, with research oscillating among one-, two-, three-, and nested-factor models without converging on a stable structure33. Similar taxonomic confusions arise in economics (“poverty”), medicine (“autism”), and computer science (“artificial intelligence”), but they are especially pernicious in the behavioral sciences. Aggregating findings across such disparate conceptualizations yields statistically significant but scientifically questionable results.
Jangle problems extend to measurement proliferation. A large-scale analysis of APA databases shows that thousands of new measures are published annually, yet over 70% are used no more than once, so the literature grows more fragmented over time37. This proliferation creates redundant research silos: “grit,” for example, shares most of its reliable variance with conscientiousness52; “psychological capital” often repackages existing well-being measures51. Semantic embeddings of construct labels can reveal more parsimonious groupings: in one analysis, the 277 distinct labels in the International Personality Item Pool were grouped into 68 embedding-based clusters38.
A third failure compounds these problems: contextual blindness conceals effect heterogeneity53. In psychology and behavioral science specifically, the most consequential omissions are demographic—age, sex, race/ethnicity, and socioeconomic status are routinely underreported8, making it impossible to determine for whom effects actually hold54. Findings robust in U.S. college samples may shrink or reverse in other cultural contexts, yet current reporting practices render such moderation invisible until replication failures accumulate. This is not merely a representational concern but a threat to validity: unreported moderators masquerade as statistical noise, obscuring the very patterns researchers seek to understand. The same logic extends beyond demographics—to firm size and industry in economics; school resources and teacher characteristics in education; dose, provider, and fidelity in intervention research; and stimuli and tasks across experimental psychology7. The general requirement is documentation of the variables over which a claim is intended to generalize.
A fourth failure—data inaccessibility and impoverishment—renders many findings functionally unverifiable and unsynthesizable. Research data frequently remain unavailable upon request9,55-58, and the trial-level granularity that verification and harmonization require is rarely present even when data are shared59. This blocks the more robust forms of evidence synthesis that depend on individual participant data (IPD) to harmonize outcomes across studies, verify or reconstruct intention-to-treat analyses, and examine effect heterogeneity across participant characteristics. The very recognition of IPD meta-analysis as a distinct and superior synthesis approach59—one that often requires manual collection of raw data from original authors—indicts the standard scientific paper’s failure as a knowledge-delivery mechanism.
These four failures compound. A study of emotion regulation conducted with one picture set in one population, reporting only summary statistics, cannot be reconciled with another that differed along all three dimensions—because the structured information required to disentangle them is precisely what the standard publication format fails to capture. Systematic reviews can still narrate such differences, but they are too slow to keep pace with literature growth, prone to error, and often yield inconclusive findings from incommensurable studies60. At the scale of more than 70,000 documented measures in psychology37 (and growing), manual curation cannot reconstitute what the format fails to record.
Generative AI promises to automate knowledge synthesis at scale, yet this potential remains largely unrealized. LLMs now permeate all stages of the writing process—with at least 13.5% of 2024 PubMed-indexed biomedical abstracts bearing AI-linked style markers61 and 22.5% of sentences in arXiv computer science abstracts estimated to be LLM-modified by September 202462. Yet they cannot synthesize knowledge from fragmented, incomparable studies. Worse, LLMs’ documented tendency to hallucinate citations and perpetuate training biases63 means that any AI-assisted synthesis must be grounded in verified, structured knowledge rather than unvetted prose.
A dual-audience architecture for the AI era
The remedy is a dual-audience article: a front-loaded narrative for human readers, backed by research-object packages for AI systems and human reanalysis (Fig. 1). The two audiences engage the same artifacts differently: machines parse code, schemas, ontology mappings, and trial-level tables, while any reader can inspect them directly—their semantics are explicit and their behavior executable, rather than asserted in prose that cannot be run. Writing for AI systems is therefore not to route content past human inspection but a way to expose structure that prose tends to obscure. Authors carry the same accountability for every component of the package as for the narrative. The same obligation extends to readers: as generative AI enters reading, review, and synthesis, expertise must be recalibrated rather than bypassed—researchers need the skills to direct AI systems, discern errors in their outputs, and check machine-generated summaries or analyses against domain standards64,65.

Fig. 1. Publication crisis and layered solution for human and AI readers. (A) Structural problems in behavioral-science publishing: stimulus fragmentation, measurement jingle–jangle, contextual (e.g., demographic) blindness, and inaccessible or impoverished data. (B) Human-optimized layers: a front-loaded narrative and interactive knowledge layer that give readers rapid access to key findings, methods, and materials. (C) Machine-optimized layer: a research-object package—executable code, semantic annotations, and tidy trial-level data—that powers interactive reanalysis for human readers and automated reproducibility checks, construct audits, and living meta-analysis for AI agents.
Inverting the narrative. A front‑loaded paper inverts traditional structure: instead of wading through a literature review before revealing findings, it begins with explicitly situated answers. The opening paragraphs should answer, in order—What did you discover? Why does it matter? How does it change our understanding?—and specify the evidence type (confirmatory, exploratory, descriptive, simulation-based, or causal), the population and setting in which the claim holds, the moderators or assumptions most likely to overturn it, and direct links to the scripts, containers, data tables, and robustness checks that reproduce or probe the headline numbers. Front-loading does not substitute summary for substance; it is the top layer of a drill-down architecture—claim, evidence, scope, methods, code, data—through which readers can probe to whatever depth their question demands.
Papers as queryable research environments. Front-loading must be paired with a citable research-object package. Each package comprises three components: an executable environment with code, containers, and dependencies; semantic annotations that link materials, measures, and constructs to persistent identifiers and shared ontologies; and tidy individual-participant-level, trial-level, and event-level data. Provenance metadata binds these components to the article and to one another, making each object independently citable and the package as a whole auditable. Any AI-assisted construction—code drafting, ontology mapping, data annotation—is logged in the same provenance layer with model and prompt versions, so that authors and reviewers can distinguish what machines drafted from what humans verified. Machines have no authorship standing; responsibility for every component rests with the humans who verify and submit it66.
Each package is built around computational reproducibility: researchers provide containerized analysis environments (e.g., Docker67 or Singularity68) that package code, dependencies, and configurations so that anyone can reproduce the analytic pipeline from raw data to final figures. Such containers are already routine in several computational fields69,70.
On this foundation sit three layers of semantic structure. First, every stimulus—whether drawn from existing databases or created anew—would receive a persistent identifier with standardized metadata for modality, normative ratings (valence, arousal, dominance), cultural validation samples, and licensing. This mirrors the Resource Identification Initiative, which assigns persistent identifiers to biological reagents and software to improve tracking and identifiability31.
Researchers creating novel stimuli document them in the same framework, contributing to an expanding queryable ecosystem. A study might reuse validated stimuli—for example, “all positive faces with arousal ratings > 7 validated in East Asian samples”—or introduce new ones such as “custom workplace scenarios rated for stress and cultural relevance”; in both cases, structured metadata enables future discovery and comparison. In language‑cognition experiments, the HED LANG framework already provides a standardized, machine‑readable vocabulary for annotating stimuli71.
Second, measurement instruments are mapped to shared conceptual spaces using ontological systems, from controlled vocabularies to formal logic-based, machine-readable ontologies72. When a study uses the Beck Depression Inventory-II, for example, individual items can be mapped to standardized clinical observation codes such as LOINC73, and the symptom dimensions they tap—negative affect, anhedonia, somatic complaints—can be linked to constructs in clinical ontologies such as the NIMH Research Domain Criteria (RDoC)74 via persistent identifiers. This semantic mapping enables AI systems to detect when ostensibly different measures assess the same construct, or when identical labels mask different phenomena, addressing the jingle–jangle problem algorithmically rather than through laborious manual coding (Box 1).
This framework makes explicit four measurement questions: What construct does this instrument measure? Why was it chosen? How are responses quantified? What study‑specific modifications were made?5 Infrastructure alone is insufficient; meaningful interoperability requires communities to forge consensus on shared definitions and standards75. The Human Behaviour Ontology illustrates this approach, systematically formalizing a controlled vocabulary of behavioral classes under shared upper-level categories to impose coherence on fragmented research domains76.
Third, data should use tidy trial-level formats linking each response to its stimulus, participant characteristics, and trial conditions, with standardized and inclusive demographic and contextual coding. This extends FAIR practice14 from dataset access to maximally reusable observations30, using community standards such as Psych-DS for behavioral data77 and BIDS for neuroimaging78.
This restructuring enables queries that are impossible with summary statistics alone— “Show effects for women over 60,” “Exclude WEIRD-dominated samples,” “How sensitive are results to different stimuli or preprocessing decisions?”—as computational operations within the executable environment, with uncertainty estimates, power warnings, and explicit flags when a query moves from confirmatory to exploratory territory. Multiverse analysis—the systematic exploration of how results vary across reasonable data-processing and analytical choices79— becomes a native capability rather than a reporting burden: alternative scripts live inside the package, the package becomes a working surface, and readers can probe a finding’s stability or build on existing work without first deciphering the author’s narrative.
Adoption is staged:
- Tier 0: Data, code, and provenance. Authors share the analytic dataset and scripts used to generate the main results in a stable repository, with persistent identifiers, clear licenses, and a brief provenance note describing recruitment, inclusion criteria, and key preprocessing steps.
- Tier 1: Tidy trial-level data and basic metadata. Shared data are restructured into a trial- or observation-level tidy format and accompanied by a simple machine-readable schema (e.g., JSON or YAML) that defines variable names, units, coding, and links between stimuli, participants, and conditions.
- Tier 2: Executable environment and automated checks. The analysis is wrapped in a containerized environment (e.g., Docker or Singularity) together with a lightweight continuous-integration script that reruns the main analyses and regenerates figures whenever the code or data change, flagging breakage early.
- Tier 3: Ontology linking and evidence-network integration. Measures, tasks, and stimuli are linked to shared ontologies and persistent identifiers; effect estimates and study-level metadata are exported in standardized form suitable for automatic ingestion into domain-specific repositories and living meta-analyses.
A minimum viable dual-audience paper corresponds to Tiers 0–1; Tiers 2–3 realize the full framework. Journals and funders should normalize Tiers 0–1 first and reserve the higher tiers for consortia, well-resourced teams, or papers whose claims depend on complex computation. Tiers 0–1 use infrastructure that is technically mature and common in some fields but unevenly normalized in behavioral science; Tier 2 is technically demanding but achievable at the lab level using widely available containerization (Docker, Singularity) and continuous-integration tooling (e.g., GitHub Actions); Tier 3 requires community infrastructure no individual lab can supply alone. The lower tiers themselves are technically available but not institutionally normalized80: a 2022 audit of empirical psychology articles found immediately accessible raw data in 14% of cases and analysis scripts in 8.5%81, with time costs, limited training, privacy exposure, uncertain standards, and absent credit for curation as the main constraints9,82,83. The framework does not require universal public release of raw data. Instead, each article should expose the most reusable package compatible with ethical, legal, and practical constraints: open code and metadata at minimum, plus an auditable data-access model ranging from synthetic or de-identified demonstration data, through controlled-access repositories, to remote-execution interfaces in which reviewers can run code against protected data without downloading them.
The relevant contrast is undocumented versus auditable, not open versus closed; this graduated logic is consistent with the modular standards of the TOP Guidelines84. Under the proposed tier system, Tier 0–1 compliance can therefore be met through any of these access models, provided two conditions are distinguished and both are met. The first is public demonstrability: a runnable pipeline against synthetic or surrogate data that lets any reader confirm the analysis executes. The second is auditable access to the actual analytic data, through controlled-access repositories, trusted research environments, or remote execution that lets an authorized reviewer confirm the reported estimates.
Box 1 | Automating the Jingle–Jangle Audit
The jingle fallacy conflates distinct phenomena under identical labels; the jangle fallacy fragments identical constructs across different names. The framework embeds a dual-layer automated audit to detect both.
The first layer employs semantic and ontological analysis. A jingle detector cross-references construct names against shared ontologies (e.g., Cognitive Atlas85), using graph algorithms to identify identical labels that occupy distinct semantic neighborhoods. In parallel, a jangle detector applies semantic embedding models to cluster scale items, flagging nominally different instruments whose item content shows high semantic convergence38,86.
The second layer provides empirical validation via statistical analysis34 of the package’s trial-level data and the living evidence network. For jangle detection, it tests extrinsic convergent validity by comparing correlation patterns between putatively different measures and external criteria. Empirical redundancy is inferred via equivalence tests when differences fall within a prespecified equivalence margin, with power adequate to detect differences smaller than that margin; audit reports must publish both the margin and the achieved power. For jingle detection, it tests discriminant validity by asking whether identically labeled measures show divergent correlations with theoretically distinct criteria—systematic differences reveal that a single label masks multiple constructs.
This dual-validation approach—semantic analysis for large-scale screening, statistical comparison for empirical confirmation—builds construct hygiene into the publication infrastructure. Continuous-integration scripts could run these audits automatically, generating curator reports that flag redundancies and collisions before they propagate. Construct screening could then become routine quality control (subject to human adjudication and ongoing validation) rather than an occasional manual exercise.
The audits screen for construct redundancy; they do not validate construct identity, which remains a theoretical judgment. They should themselves be evaluated for reproducibility of equivalence inferences and for potential hallucination of LLM-generated ontological links.
Curator reports therefore separate deterministic checks from LLM-generated inferences, log model and prompt versions, preserve source spans and code or data references, and require human adjudication before any finding feeds back into shared ontologies or downstream syntheses.
Restructuring peer review for executable verification
The dual‑audience framework also restructures peer review. A front‑loaded narrative lets reviewers judge the contribution’s novelty and significance before parsing methods. More consequentially, the research-object package makes the methods auditable; current peer review rarely scrutinizes data and code—one large-scale attempt found that only 879 of 15,817 attempted executions of Python notebooks reproduced identical outputs, suggesting that methodological claims are largely taken on trust87. Reviewers can then inspect configuration files and data, execute code, and verify that figures regenerate from the underlying data.
The same package provides the essential infrastructure for AI-assisted review (Fig. 2).
Article text alone—even when serialized as JATS XML—is insufficient for AI systems that must verify tables, figures, data, and code. In one benchmark in Alzheimer’s disease, autonomous AI agents reproduced only about half of 35 findings across five papers when relying only on prose methods sections; ambiguous or incomplete text led them to apply incorrect or incomplete statistical methods88. Such agent-generated reconstruction is a useful fallback when no original code exists, but it is not (currently) a substitute for the executable workflow an author can package with the article. Structured, machine-friendly formats and tools (e.g., CSV, Markdown, Git) open the way to AI-driven quality control: validating references, auditing logical consistency, checking mathematical and statistical accuracy, and systematically verifying the package. Structured packages also let an LLM delegate reading to deterministic interfaces and reserve generation for synthesis and explanation; an analogous division has been shown to outperform direct text serialization of structured content in multi-step reasoning39. Verification then moves beyond confirming that code runs to running automated multiverse analyses that vary data‑processing choices and analytic parameters to map the fragility of a study’s conclusions.
Formal ontologies shift the reviewer’s task: rather than excavating construct definitions from inconsistent terminology, reviewers can see where definitions agree, diverge, or require adjudication89. Human reviewers receive both the manuscript and an AI‑generated audit report and can direct their attention to interpretive claims, novelty, and theoretical significance— judgments that require domain expertise and that automated checks cannot make.

Fig. 2. Hybrid AI–human peer review enabled by research-object packages. Authors submit a front‑loaded manuscript paired with its research-object package (code, data, containers, ontologies). AI review agents check reproducibility, multiverse‑robustness, and citation integrity to generate an audit summarizing status, warnings, and fragility. Human reviewers and editors use this audit alongside the narrative to request targeted author responses, assess conceptual novelty and theoretical contribution, adjudicate interpretation versus evidence, and make final publication decisions.
Constructing living evidence networks
Beyond improving communication and quality assurance, the dual-audience framework creates the technical foundation for systematic knowledge integration. Papers can be linked into living evidence networks—extensions of the manually intensive “living systematic review” model90—in which each publication becomes a node that automatically contributes to evolving synthesis. Such networks have two coupled layers: a knowledge layer, in which article text, claims, citations, methods, and ontological mappings support retrieval-augmented or graph-based queries over what has been claimed, disputed, replicated, or retracted; and a synthesis layer, in which effect estimates, IPD summaries, contextual moderators, preregistration status, data and code versions, and quality indicators feed living meta-analyses.
Consider how systematic reviews and meta-analyses currently work: researchers manually search databases, screen thousands of abstracts, extract data from hundreds of PDFs, code effect sizes, and produce static summary estimates—a process that often takes more than a year and involves several authors91. Yet by the time of publication, many meta-analyses are already outdated: median “shelf life” is 5.5 years, with 23% requiring updates within two years and 7% already in need of updating at the time of publication92. New studies accumulate in limbo until another large manual effort is mounted. This process is further compromised by systematic publication bias: null or inconvenient results remain buried in file drawers while false positives proliferate93. Fig. 3 contrasts this static, labor-intensive pipeline with the living evidence network enabled when structured papers feed domain-specific repositories and continuously updated meta-analyses.
Living evidence networks invert this workflow. When researchers publish under this framework—whether in journals or public registries94—standardized exports from their research-object packages could populate theme-specific evidence repositories. A mindfulness–anxiety study, for example, would contribute its effect sizes, sample and context variables, and methodological features to the running meta-analysis regardless of whether results are significant, null, or contradictory. Pooled estimates would refresh through rolling, versioned releases; retractions would be flagged in the next release rather than after years of latency, automating and accelerating the update process currently managed by teams conducting living systematic reviews95.
The same structured knowledge base could also become training data for AI models that predict outcomes for novel interventions or specific populations60. With validation against held-out populations and provided that causal-identification conditions, estimands, and transportability assumptions are explicitly encoded, such models could in principle support forecasts of intervention efficacy for a given demographic—a use of evidence synthesis as actionable guidance rather than only retrospective summary. Data rights, privacy, consent, and the terms of model training and retention must nonetheless be specified explicitly, since the same structures that enable synthesis can be misused if these governance conditions are not met.

Fig. 3. From static meta-analyses to living evidence networks. (A) Traditional workflow in which scattered PDFs feed a one-off meta-analysis requiring months or years of manual screening and coding, yielding a synthesis that is already outdated at publication. (B) Articles and their research-object packages contribute effect sizes, moderators, and quality indicators into domain-specific repositories, and repository-curated updates—new estimates, retraction notices—flow back to each article, so that information moves in both directions rather than only from article to repository. (C) Living evidence network in which domain-specific repositories continuously update living meta-analyses and cross-domain syntheses; new studies, retractions, and registered null results dynamically alter study weights, providing rapidly updated evidence for clinical guidelines, policy briefs, and AI prediction models.
This framework requires two additional infrastructure components beyond the semantic anchoring and ontological mapping already described. First, research artifacts—datasets, code, preregistrations, and derived effect-size tables—must be treated as versioned, citable objects.
Version-control systems such as Git provide transparent audit trails, but propagating corrections requires additional layers: persistent identifiers, standardized cross‑reference metadata (e.g., DataCite‑style schemas), explicit dependency graphs linking studies to syntheses, and registries that index these links. When a dataset or analysis is corrected or retracted, systems built on this graph can automatically flag and update downstream meta‑analyses and evidence syntheses, instead of relying on formal errata and retractions that are slow and cumbersome to issue96 and often remain invisible for years97.
Second, quality indicators—sample size, preregistration status, replication attempts, methodological rigor scores—should drive stratified analyses, sensitivity analyses excluding high-risk studies, bias-adjustment models, and transparent evidence grading (e.g., GRADE), rather than formal quality weights (which Cochrane has cautioned against absent an empirical basis for the choice of weights). Standard precision weighting (inverse-variance) remains the default; study quality enters as a moderator in meta-regression and as the basis for stratified summaries, so that evidence synthesis reflects both the quantity and quality of available research.
Maintaining living evidence networks requires federated stewardship. Journals provide article-object metadata; repositories host versioned executable research objects; domain societies or designated registry boards maintain construct and effect-size records; curators adjudicate contested mappings and retractions; and aggregators index typed links among articles, data, code, and synthesis nodes. This indexing function is distinct from general discovery aggregators such as Google Scholar or OpenAlex, which do not generally expose full typed provenance and retraction metadata. Updates should propagate through versioned releases with auditable curator logs rather than silent mutation. Each network requires a designated domain steward—a registry board, society working group, or repository-hosted curation team—empowered to issue versioned releases, adjudicate mappings, and log decisions with revision history. Ultimately, curator authority remains a governance challenge, and it should be interoperable with established infrastructures such as Crossref and ORCID and accountable to domain-specific stewardship bodies.
Navigating implementation risks and inequities
The framework outlined above is aspirational; without attention to implementation constraints, it could reinforce existing inequities. Building executable containers, tidy trial‑level datasets, and ontological mappings requires technical capacity that many small labs, non‑elite institutions, and practitioner settings do not yet have. If research-object packages become de facto requirements without parallel investments in infrastructure, training, and credit, adoption will be slow and skewed toward well‑resourced groups. Standardization of shared ontologies and demographic vocabularies also entails substantial coordination costs and risks freezing contested constructs or marginalizing alternative theoretical traditions unless governance is explicitly pluralistic, revisable, and transparent.
Sensitive and confidential data pose additional risks. In clinical, educational, and small or marginalized populations, naive mandates for fully open trial‑level data collide with privacy protections, data‑sovereignty claims, and legal or ethical constraints. The dual‑audience architecture must therefore support tiered access, secure data enclaves, and remote‑execution or synthetic‑data solutions so that code and metadata remain reusable even when raw data cannot be widely shared. Practically, the package exposes auditable materials—protocols, code, schemas, data dictionaries, validation reports, synthetic demonstration data, controlled-access procedures, and remote-execution endpoints—without exposing what must remain protected. Synthetic data are labeled as demonstration data rather than as substitutes for confirmatory analysis unless their inferential validity has been evaluated.
Finally, lowering the friction for reanalysis also lowers the friction for motivated misuse.
Interactive reanalysis interfaces can make it easier for ideological actors to cherry‑pick specifications, ignore multiverse fragility or quality indicators, and promote “do your own research” narratives that overstate the certainty of convenient results. Design choices can mitigate these risks by foregrounding robustness summaries rather than single estimates, making departures from preregistered analyses and default pipelines explicitly visible, and tying reanalyses back into version‑controlled living evidence networks where idiosyncratic claims are evaluated against the full corpus rather than in isolation. A related risk is the proliferation of formulaic secondary literatures. AI-ready public datasets can be mined into single-factor association papers that ignore interactions, choose subsets selectively, skip multiple-testing correction, and—at the extreme—feed paper-mill production lines; a recent analysis of NHANES-derived publications documents this pattern at scale98. The remedy is not to make datasets unusable but to make reuse auditable: preregistration of confirmatory reuse, principled justification of subset selection with multiple-testing correction, reuse identifiers for high-value datasets, and editorial screening for formulaic designs. Living evidence networks should label exploratory reuse separately from confirmatory evidence and weight syntheses by design quality rather than publication count.
A roadmap for federated stewardship
Technical standardization and platform coordination. Publishers, scholarly societies, repositories, and research libraries all face a strategic choice: continue distributing static PDFs as open‑access mandates compress subscription revenues, or adopt and converge on open, shared formats for structured, executable packages that any interface can use—and help build the infrastructure that makes those papers queryable and verifiable.
For any hosting platform, two implementation paths emerge: develop native AI systems—tools trained on scientific content with domain expertise in methodology, statistics, and interpretation—or leverage browser‑ or operating‑system–level AI. If publishers expose structured paper contexts via open APIs, tools such as Gemini in Chrome could access the full research-object package rather than only rendered text. Either approach enables text‑based services such as on‑demand translation that preserves technical precision, personalized summarization tailored to reader expertise99, and advanced Q&A that can execute code for reanalysis or generate new visualizations. Audio and video services could automatically produce conversational podcasts or video summaries, making research accessible across formats and languages.
To remain viable as AI systems and interfaces evolve, these implementations should rest on open, versioned standards. The archival “paper package”—identifiers, metadata, data, and executable code—must remain portable across hosts and decoupled from any single AI model or user interface, so that the scientific record outlives particular vendors and tools. Publishers should also treat machine-readable serializations of the article itself—JATS XML and Markdown alongside the rendered PDF—as standard deliverables rather than typesetting byproducts; this relatively low-cost change removes the parsing layer on which most AI ingestion errors are concentrated. They should likewise expose typed bidirectional links between each article and the objects in its package—datasets, code, stimuli, protocols, preregistrations— using existing Crossref and DataCite relationship metadata so that the connections are visible to both readers and machines.
These services imply a shift from charging primarily for document access toward supporting authenticated analytical capability. Funding models may range from institutional support and public infrastructure to subscription‑based access to advanced tools, while the underlying narrative text and research-object packages remain findable and portable even when specific interfaces are restricted. For sensitive data, access would operate through institutional agreements that create a trusted “data commons” of authorized researchers.
Proofs of concept already exist. Table 1 maps deployed infrastructures and packaging standards against the proposed tiers, showing that each tier finds a precursor in at least one existing system, many of them in adjacent disciplines. Existing work proves feasibility; the framework specifies the missing layer—broadly adopted identifiers for stimuli, tasks, measures, and instruments; tidy trial-level schemas with demographic and contextual coding; construct ontologies that make measurement-equivalence claims explicit and auditable; and standardized exports for living evidence networks.
| Initiative (type) | Closest tier alignment | Capability demonstrated | What is missing for the proposed framework |
|---|---|---|---|
| Resource Identification Initiative / Research Resource Identifiers (RRIDs; registry)31 | Identifier substrate for Tiers 0–3 | Persistent, unambiguous identifiers for antibodies, cell lines, model organisms, plasmids, and software/tools/databases, improving traceability across papers and disciplines | No broadly adopted RRID-like identifier layer for psychological stimuli, tasks, measures, instruments, or construct mappings |
| RO-Crate (packaging standard)15 | Cross-tier packaging substrate; strongest for Tiers 0–2 | FAIR-aligned bundling of datasets, code, and provenance with rich JSON-LD metadata; interoperable with workflow systems and repositories | Schema-agnostic about meaning: provides slots but does not specify which ontologies populate them, nor whether “executive function” in one container refers to the same construct as in another |
| Frictionless Data Package (packaging standard)18 | Tier 1 substrate | Lightweight tabular packaging with machine-readable schemas for variable types, units, and constraints; low adoption cost | Optimized mainly for tabular schemas; provides no native layer for stimuli, executable environments, container provenance, or construct-to-instrument mappings |
| Croissant (ML dataset standard)100 | Tier 1 substrate for ML-ready datasets | Vocabulary for documenting machine-learning datasets— features, labels, splits, licensing—with consistent ingestion into ML toolchains | Built around ML training conventions rather than experimental psychology; no construct ontology, trial-level experimental schema, or demographic and contextual coding norms |
| Databrary (controlled-access repository)30 | Tiers 0–1 with access governance | Sharing of behavioral video and other identifiable data under researcher-credentialed access agreements; an existence proof that openness need not mean fully public | Not designed as a general standard for tidy trial-level schemas, executable workflows, or automated reproducibility checks; provides metadata and access infrastructure but not these layers |
| eLife Executable Research Articles (publishing platform)20 | Tier 1 + partial Tier 2 | Embedded data, code, and figure/table elements that readers can inspect, modify, and rerun in the browser without local installation | Article-scoped execution; no behavioral-science stimulus identifiers, construct mappings, or evidence-network exports |
| Code Ocean (compute platform)101 | Tier 2 | Containerized code, data, and dependencies as cloud-rerunnable capsules; decouples execution from local environment | Discipline-neutral execution layer; no construct ontology, demographic/contextual coding, or downstream synthesis exports |
| EBRAINS Live Papers (publishing platform)23 | Tiers 1–2 with neuroscience-specific, Tier-3-adjacent provenance | Code, data, models, and interactive simulation/visualization bundled within neuroscience publications | Domain-specific to computational neuroscience; no general behavioral-science layer for construct audits, trial-level psychological data, or standardized meta-analytic exports |
| Nanopublications (semantic web standard)21,22 | Tier 3 substrate; evidence-network layer | Atomic scientific assertions—claim, provenance, publication provenance—encoded as citable RDF entities with persistent Trusty URIs (Uniform Resource Identifiers); decentralized server network supporting millions of published nanopublications across domains | Scope limited to single assertions rather than full research-object bundles; no construct ontology, stimulus or task identifier integration, or trial-level schema; adoption in psychology negligible |
| Schol-AR (presentation layer)102 | Interactive presentation layer; complements Tiers 1–2 | Manipulable visualizations and AR-embedded data displays inside web or print article views | Display layer only; no execution, semantic mapping, provenance, or evidence-network ingestion |
Note. Most initiatives listed here supply components of a tier or provide substrate on which one could be built; “closest tier alignment” denotes correspondence, not full implementation.
Whether the proposed structure pays for itself is an empirical question. Comparing AI extraction accuracy, citation accuracy, table recovery, reproduction success, and hallucination rate across PDF-only, machine-readable text, and full structured-package conditions would establish where the marginal benefit justifies the marginal author cost. Feasibility should be reported alongside performance: author preparation time per tier, curator labor per submission, code-execution success rate (distinct from full reproduction success), frequency of privacy-driven exceptions, reviewer burden, and downstream reuse.
The next evolution demands clear governance as well as technical innovation. Three questions are central. First, platforms hosting AI–paper interactions must protect the privacy of user queries and interaction logs. Second, they must distinguish policies for confidential peer-review materials from those for published content, and avoid sending unpublished manuscripts to external commercial AI systems that may retain proprietary data66. Third, publisher AI policies should disaggregate uses often collapsed under “LLM ingestion”: retrieval or indexing of the published article; retrieval-augmented generation over the article and its research-object package; model training or fine-tuning; and retention or analysis of reader, author, and reviewer interaction logs. These uses differ legally and ethically: open-access text may be searchable under its license while still raising attribution and rights-reservation questions for training; packages may contain personal or culturally sensitive material even when the article is public; peer-review materials remain privileged communications. Publishers should disclose at submission and publication which uses are permitted, which are opt-in or opt-out, what is retained, whether third-party vendors receive content, and whether interaction logs feed product development or model training. Authors should retain rights-reservation options consistent with applicable law, and the underlying architectures must guarantee privacy-first design: user interactions encrypted, data uses—including for training—transparent and subject to meaningful consent and oversight.
Institutional incentives and career reform. Universities and research institutions control the most powerful lever for adoption: career incentives. Promotion and tenure committees have historically undervalued digital research assets, creating a systemic disincentive to produce the high-quality data and code essential for a cumulative science103.
Institutions should revise hiring, promotion, and tenure criteria to credit the contributions a research-object package makes visible: curated datasets, executable code, validated stimuli and measures, ontological mappings, and the systematic consensus-building that yields shared terminologies and methodological standards75. This aligns with the Coalition for Advancing Research Assessment (CoARA)104 and the Declaration on Research Assessment (DORA)105, which call for assessment of diverse outputs beyond journal impact factors, and with the CRediT taxonomy106, which already provides standardized roles—data curation, software, resources, validation, visualization—through which contributors receive explicit credit for the work of creating those objects, not only for co-authorship of the article. Beyond explicit credit, citation and reuse metrics for data and software render object-level contributions measurable, so that assessment systems can weight evidence of impact at the object level rather than article authorship alone.
Recognition alone is insufficient. Institutions must provide funding, computational resources, version-control systems, and technical training that make comprehensive documentation feasible rather than burdensome. Research libraries should expand from literature access to data curation, helping faculty turn messy research outputs into structured, queryable packages. Early adopters may gain visibility and reuse advantages, consistent with evidence that openly shared datasets attract higher citation rates80. As AI-driven discovery tools emerge, institutions producing structured, machine-readable research will see their faculty become more visible and influential, creating systematic advantages in knowledge synthesis and collaboration.
Conclusions and outlook
This transformation shifts authority from uncheckable narration to contestable evidence: readers have traditionally had to trust the narrative because data, code, and analytic choices were inaccessible; the research-object package makes those materials inspectable and executable. AI systems make evidence easier to query, but their outputs remain interpretations requiring human judgment, theoretical context, and accountability. Reanalysis of a persistent, executable research object moves beyond static replication toward dynamic exploration, making robustness checks, alternative specifications, and hypothesis generation cheap. Such reanalyses remain exploratory: unbiased confirmatory tests still require prespecified—ideally preregistered—design and analyses, and evaluation on fresh or held‑out data not used to generate the hypotheses or make the analytic decisions.
The scaffold is discipline‑agnostic. Psychology is the stress test because its core objects—stimulus sets, tasks, and jingle–jangle‑prone constructs—are unusually tangled; other domains can substitute their own reagents, instruments, specimens, datasets—or, in interpretive fields, corpora, editions, translations, and annotation provenance. Widespread adoption would generate structured research artifacts that surpass the unstructured text currently used to train most models, giving AI systems better-structured inputs on which statistical engines and verification tools can operate—and giving scholarly communities the standing to set those tools’ defaults rather than inherit them.
References
- Bornmann, L., Haunschild, R. & Mutz, R. Growth rates of modern science: a latent piecewise growth curve approach to model publication numbers from established and new literature Humanities and Social Sciences Communications 8, 224 (2021). https://doi.org/10.1057/s41599-021-00903-w
- National Science Publications output: U.S. trends and international comparisons. (National Science Foundation, 2024).
- Tenopir, C., King, D. W., Christian, L. & Volentine, R. Scholarly article seeking, reading, and use: A continuing evolution from print to electronic in the sciences and social Learned Publishing 28, 93-105 (2015). https://doi.org/10.1087/20150203
- Lin, Z. Why and how to embrace AI such as ChatGPT in your academic life. R. Soc. Open Sci. 10, 230658 (2023). https://doi.org/10.1098/rsos.230658
- Flake, K. & Fried, E. I. Measurement schmeasurement: Questionable measurement practices and how to avoid them. Adv. Meth. Pract. Psychol. Sci. 3, 456-465 (2020). https://doi.org/10.1177/2515245920952393
- Clark, H. H. The language-as-fixed-effect fallacy: A critique of language statistics in psychological Journal of Verbal Learning and Verbal Behavior 12, 335-359 (1973). https://doi.org/10.1016/S0022-5371(73)80014-3
- Yarkoni, T. The generalizability crisis. Behav. Brain Sci. 45, e1 (2022). https://doi.org/10.1017/s0140525x20001685
- Sterling, E., Pearl, H., Liu, Z., Allen, J. W. & Fleischer, C. C. Demographic reporting across a decade of neuroimaging: A systematic Brain Imaging and Behavior 16, 2785-2796 (2022). https://doi.org/10.1007/s11682-022-00724-8
- Tedersoo, L. et al. Data sharing practices and data availability upon request differ across scientific disciplines. Scientific Data 8, 192 (2021). https://doi.org/10.1038/s41597-021–00981-0
- Hardwicke, T. E. et al. Estimating the prevalence of transparency and reproducibility-related research practices in psychology (2014–2017). Psychol. Sci. 17, 239-251 (2022). https://doi.org/10.1177/1745691620979806
- National Academies of Sciences, Engineering, and Medicine. Ontologies in the Behavioral Sciences: Accelerating Research and the Spread of Knowledge. (National Academies Press, 2022).
- Lála, J. et al. PaperQA: Retrieval-augmented generative agent for scientific research. arXiv:2312.07559 (2023). https://doi.org/10.48550/arXiv.2312.07559
- Livathinos, N. et al. in AAAI 25: Workshop on Open-Source AI for Mainstream Use (Association for the Advancement of Artificial Intelligence, Philadelphia, PA, USA, 2025).
- Wilkinson, M. D. et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3, 160018 (2016). https://doi.org/10.1038/sdata.2016.18
- Soiland-Reyes, S. et al. Packaging research artefacts with RO-Crate. Data Science 5, 97-138 (2022). https://doi.org/10.3233/DS-210053
- Bourne, E. et al. Improving the future of research communications and e-scholarship (Dagstuhl perspectives workshop 11331). Dagstuhl Manifestos 1, 41-60 (2012). https://doi.org/10.4230/DagMan.1.1.41
- Bechhofer, et al. Why linked data is not enough for scientists. Future Generation Computer Systems 29, 599-611 (2013). https://doi.org/10.1016/j.future.2011.08.004
- Fowler, D., Barratt, J. & Walsh, P. Frictionless Data: Making research data quality visible. International Journal of Digital Curation 12, 274-285 (2018). https://doi.org/10.2218/ijdc.v12i2.577
- Edmunds, C. et al. Experiences in integrated data and research object publishing using GigaDB. International Journal on Digital Libraries 18, 99-111 (2017). https://doi.org/10.1007/s00799-016-0174-6
- Maciocci, G., Aufreiter, M. & Bentley, N. Introducing eLife’s first computationally reproducible article, <https://elifesciences.org/labs/ad58f08d/introducing-elife-s-first–computationally-reproducible-article> (2019).
- Pensoft Publishers & Knowledge How it works: Nanopublications linked to articles in RIO Journal, <https://blog.pensoft.net/2023/05/17/how-it-works–nanopublications-linked-to-articles-in-rio-journal/> (2023).
- Schultes, E. et al. The comparative anatomy of nanopublications and FAIR Digital Objects. Research Ideas and Outcomes 8, e94150 (2022). https://doi.org/10.3897/rio.8.e94150
- Appukuttan, S., Bologna, L. L., Schürmann, F., Migliore, M. & Davison, A. P. EBRAINS Live Papers – Interactive resource sheets for computational studies in neuroscience. Neuroinformatics 21, 101-113 (2023). https://doi.org/10.1007/s12021-022–09598-z
- Berman, F. & Crosas, M. The Research Data Alliance: Benefits and challenges of building a community organization. Harvard Data Science Review 2 (2020). https://doi.org/10.1162/99608f92.5e126552
- GO GO FAIR Initiative, <https://www.go-fair.org/go-fair-initiative/> (Accessed May 1, 2026).
- Barcelona Declaration on Open Research Barcelona Declaration on Open Research Information, <https://barcelona-declaration.org/> (2024).
- Mullen, B. Open access, scholarly communication, and open science in psychology: An overview for researchers. Sage Open 14 (2024). https://doi.org/10.1177/21582440231205390
- Chambers, D. & Tzavella, L. The past, present and future of Registered Reports. Nat. Hum. Behav. 6, 29-42 (2022). https://doi.org/10.1038/s41562-021-01193-7
- Nosek, A., Ebersole, C. R., DeHaven, A. C. & Mellor, D. T. The preregistration revolution. Proc. Natl. Acad. Sci. U.S.A. 115, 2600-2606 (2018). https://doi.org/10.1073/pnas.1708274114
- Gilmore, R. O., Kennedy, J. L. & Adolph, K. E. Practical solutions for sharing data and materials from psychological Adv. Meth. Pract. Psychol. Sci. 1, 121-130 (2018). https://doi.org/10.1177/2515245917746500
- Bandrowski, A. et al. The Resource Identification Initiative: A cultural shift in publishing. Comp. Neurol. 524, 8-22 (2016). https://doi.org/10.1002/cne.23913
- Baggetta, & Alexander, P. A. Conceptualization and operationalization of executive function. Mind, Brain, and Education 10, 10-33 (2016). https://doi.org/10.1111/mbe.12100
- Karr, E. et al. The unity and diversity of executive functions: A systematic review and re-analysis of latent variable studies. Psychol. Bull. 144, 1147-1185 (2018). https://doi.org/10.1037/bul0000160
- Gonzalez, O., MacKinnon, D. P. & Muniz, F. B. Extrinsic convergent validity evidence to prevent jingle and jangle Multivariate Behavioral Research 56, 3-19 (2021). https://doi.org/10.1080/00273171.2019.1707061
- Diconne, K., Kountouriotis, G. K., Paltoglou, A. E., Parker, A. & Hostler, T. J. Presenting KAPODI—the searchable database of emotional stimuli Emotion Review 14, 84-95 (2022). https://doi.org/10.1177/17540739211072803
- Lin, Z., Ma, Q., Huang, X., Wu, X. & Zhang, Y. Pervasive failure to report properties of visual stimuli in experimental research in psychology and neuroscience: Two metascientific studies. Psychol. Bull. 149, 487-505 (2023). https://doi.org/10.1037/bul0000399
- Anvari, F. et al. Defragmenting psychology. Nat. Hum. Behav. 9, 836-839 (2025). https://doi.org/10.1038/s41562-025-02138-0
- Wulff, D. U. & Mata, R. Semantic embeddings reveal and address taxonomic incommensurability in psychological Nat. Hum. Behav. 9, 944-954 (2025). https://doi.org/10.1038/s41562-024-02089-y
- Jiang, et al. in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 9237-9251 (Association for Computational Linguistics).
- Hanson, A., Barreiro, P. G., Crosetto, P. & Brockington, D. The strain on scientific publishing. Quantitative Science Studies 5, 823-843 (2024). https://doi.org/10.1162/qss_a_00327
- Lang, P. J., Bradley, M. M. & Cuthbert, B. N. International affective picture system (IAPS): Affective ratings of pictures and instruction manual. (NIMH, Center for the Study of Emotion & Attention, Technical Report A-6, Gainesville, FL, 2005).
- Kurdi, B., Lozano, S. & Banaji, M. R. Introducing the Open Affective Standardized Image Set (OASIS). Behav. Res. Methods 49, 457-470 (2017). https://doi.org/10.3758/s13428-016-0715-3
- Dan-Glauser, S. & Scherer, K. R. The Geneva affective picture database (GAPED): A new 730-picture database focusing on valence and normative significance. Behav. Res. Methods 43, 468-477 (2011). https://doi.org/10.3758/s13428-011-0064-1
- Marchewka, A., Żurawski, Ł., Jednoróg, K. & Grabowska, A. The Nencki Affective Picture System (NAPS): Introduction to a novel, standardized, wide-range, high-quality, realistic picture database. Res. Methods 46, 596-610 (2014). https://doi.org/10.3758/s13428-013-0379-1
- Weierich, R., Kleshchova, O., Rieder, J. K. & Reilly, D. M. The Complex Affective Scene Set (COMPASS): Solving the social content problem in affective visual stimulus sets. Collabra: Psychology 5, 53 (2019). https://doi.org/10.1525/collabra.256
- Mancuso, V. et al. IAVRS—International Affective Virtual Reality System: Psychometric assessment of 360° images by using psychophysiological Sensors 24, 4204 (2024). https://doi.org/10.3390/s24134204
- Balsamo, M., Carlucci, L., Padulo, C., Perfetti, B. & Fairfield, B. A bottom-up validation of the IAPS, GAPED, and NAPS affective picture databases: Differential effects on behavioral performance. Front Psychol 11, 2187 (2020). https://doi.org/10.3389/fpsyg.2020.02187
- Michelson, M. & Reuter, K. The significant cost of systematic reviews and meta-analyses: A call for greater involvement of machine learning to assess the promise of clinical trials. Contemporary Clinical Trials Communications 16, 100443 (2019). https://doi.org/10.1016/j.conctc.2019.100443
- Thorndike, E. L. An Introduction to the Theory of Mental and Social Measurements. (The Science Press, 1904).
- Kelley, T. L. Interpretation of educational measurements. (World Book Company, 1927).
- van Zyl, E. & Rothmann, S. Grand challenges for positive psychology: Future perspectives and opportunities. Front Psychol 13, 833057 (2022). https://doi.org/10.3389/fpsyg.2022.833057
- Ponnock, A. et al. Grit and conscientiousness: Another jangle fallacy. J Res Pers 89, 104021 (2020). https://doi.org/10.1016/j.jrp.2020.104021
- von Hippel, T. & Schuetze, B. A. How not to fool ourselves about heterogeneity of treatment effects. Adv. Meth. Pract. Psychol. Sci. 8, 25152459241304347 (2025). https://doi.org/10.1177/25152459241304347
- Call, C. C. et al. An ethics and social-justice approach to collecting and using demographic data for psychological Perspect. Psychol. Sci. 18, 979-995 (2023). https://doi.org/10.1177/17456916221137350
- Danchev, V., Min, Y., Borghi, J., Baiocchi, M. & Ioannidis, J. P. A. Evaluation of data sharing after implementation of the International Committee of Medical Journal Editors data sharing statement requirement. JAMA Network Open 4, e2033972 (2021). https://doi.org/10.1001/jamanetworkopen.2020.33972
- Wicherts, M., Borsboom, D., Kats, J. & Molenaar, D. The poor availability of psychological research data for reanalysis. Am. Psychol. 61, 726-728 (2006). https://doi.org/10.1037/0003-066X.61.7.726
- Vines, H. et al. The availability of research data declines rapidly with article age. Curr. Biol. 24, 94-97 (2014). https://doi.org/10.1016/j.cub.2013.11.014
- Hardwicke, T. E. & Ioannidis, J. P. A. Populating the Data Ark: An attempt to retrieve, preserve, and liberate data from the most highly-cited psychology and psychiatry PLOS ONE 13, e0201856 (2018). https://doi.org/10.1371/journal.pone.0201856
- Tierney, J. F., Stewart, L. A., Clarke, M. & on behalf of the Cochrane Individual Participant Data Meta-analysis Methods in Cochrane Handbook for Systematic Reviews of Interventions 643-658 (Wiley, 2019).
- Castro, O., Mair, J., von Wangenheim, F. & Kowatsch, T. in Proceedings of the 17th International Joint Conference on Biomedical Engineering Systems and Technologies -HEALTHINF. 671-678 (SciTePress).
- Kobak, , González-Márquez, R., Horvát, E.-Á. & Lause, J. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances 11, eadt3813 (2025). https://doi.org/10.1126/sciadv.adt3813
- Liang, W. et al. Quantifying large language model usage in scientific papers. Nat. Hum. Behav. 9, 2599-2609 (2025). https://doi.org/10.1038/s41562-025-02273-8
- Walters, W. H. & Wilder, E. I. Fabrication and errors in the bibliographic citations generated by Sci. Rep. 13, 14045 (2023). https://doi.org/10.1038/s41598-023–41032-5
- Lin, & Sohail, A. Recalibrating academic expertise in the age of generative AI. Patterns 7, 101473 (2026). https://doi.org/10.1016/j.patter.2025.101473
- Resnik, B., Hosseini, M. & Hauswald, R. Autonomous artificial intelligence, scientific research, and human values. AI Ethics 6, 141 (2026). https://doi.org/10.1007/s43681-025–00908-0
- Lin, Towards an AI policy framework in scholarly publishing. Trends Cogn. Sci. 28, 85-88 (2024). https://doi.org/10.1016/j.tics.2023.12.002
- Boettiger, C. An introduction to Docker for reproducible research. SIGOPS Oper. Syst. Rev. 49, 71–79 (2015). https://doi.org/10.1145/2723872.2723882
- Kurtzer, M., Sochat, V. & Bauer, M. W. Singularity: Scientific containers for mobility of compute. PLOS ONE 12, e0177459 (2017). https://doi.org/10.1371/journal.pone.0177459
- Moreau, D., Wiebels, K. & Boettiger, C. Containers for computational reproducibility. Nature Reviews Methods Primers 3, 50 (2023). https://doi.org/10.1038/s43586-023–00236-9
- Di Tommaso, et al. Nextflow enables reproducible computational workflows. Nat. Biotechnol. 35, 316-319 (2017). https://doi.org/10.1038/nbt.3820
- Denissen, M., Pöll, B., Robbins, K., Makeig, S. & Hutzler, F. HED LANG – A Hierarchical Event Descriptors library extension for annotation of language cognition experiments. Scientific Data 11, 1428 (2024). https://doi.org/10.1038/s41597-024-04282–0
- Sharp, C., Kaplan, R. M. & Strauman, T. J. The use of ontologies to accelerate the behavioral sciences: Promises and Curr. Dir. Psychol. Sci. 32, 418-426 (2023). https://doi.org/10.1177/09637214231183917
- McDonald, J. et al. LOINC, a universal standard for identifying laboratory observations: A 5-year update. Clin. Chem. 49, 624-633 (2003). https://doi.org/10.1373/49.4.624
- Insel, et al. Research domain criteria (RDoC): Toward a new classification framework for research on mental disorders. Am. J. Psychiatry 167, 748-751 (2010). https://doi.org/10.1176/appi.ajp.2010.09091379
- Leising, D., Liesefeld, H., Buecker, S., Glöckner, A. & Lortsch, S. A tentative roadmap for consensus building processes. Personality Science 5, 27000710241298610 (2024). https://doi.org/10.1177/27000710241298610
- Schenk, P. et al. An ontological framework for organising and describing behaviours: The Human Behaviour Ontology [version 2; peer review: 1 approved, 2 approved with reservations]. Wellcome Open Research 9, 237 (2025). https://doi.org/10.12688/wellcomeopenres.21252.2
- Psych-DS Working Group. Psych-DS: A simple, easy-to-adopt format for sharing behavioral and social science data, <https://psychds-docs.readthedocs.io/en/latest/> (Accessed May 1, 2026).
- Gorgolewski, J. et al. The brain imaging data structure, a format for organizing and describing outputs of neuroimaging experiments. Scientific Data 3, 160044 (2016). https://doi.org/10.1038/sdata.2016.44
- Steegen, S., Tuerlinckx, F., Gelman, A. & Vanpaemel, W. Increasing transparency through a multiverse analysis. Perspect. Psychol. Sci. 11, 702-712 (2016). https://doi.org/10.1177/1745691616658637
- Piwowar, A. & Vision, T. J. Data reuse and the open data citation advantage. PeerJ 1, e175 (2013). https://doi.org/10.7717/peerj.175
- Hardwicke, T. E. et al. Prevalence of transparent research practices in psychology: A cross-sectional study of empirical articles published in Adv. Meth. Pract. Psychol. Sci. 7, 25152459241283477 (2024). https://doi.org/10.1177/25152459241283477
- Houtkoop, L. et al. Data sharing in psychology: A survey on barriers and preconditions. Adv. Meth. Pract. Psychol. Sci. 1, 70-85 (2018). https://doi.org/10.1177/2515245917751886
- Borycz, J. et al. Perceived benefits of open data are improving but scientists still lack resources, skills, and Humanities and Social Sciences Communications 10, 339 (2023). https://doi.org/10.1057/s41599-023-01831-7
- Nosek, A. et al. Promoting an open research culture. Science 348, 1422-1425 (2015). https://doi.org/10.1126/science.aab2374
- Poldrack, A. et al. The Cognitive Atlas: Toward a knowledge foundation for cognitive neuroscience. Frontiers in Neuroinformatics 5, 17 (2011). https://doi.org/10.3389/fninf.2011.00017
- Huang, Z., Long, Y., Peng, K. & Tong, S. An embedding-based semantic analysis approach: A preliminary study on redundancy detection in psychological concepts operationalized by scales. Journal of Intelligence 13, 11 (2025). https://doi.org/10.3390/jintelligence13010011
- Samuel, & Mietchen, D. Computational reproducibility of Jupyter notebooks from biomedical publications. GigaScience 13, giad113 (2024). https://doi.org/10.1093/gigascience/giad113
- Dobbins, , Xiong, C., Lan, K. & Yetisgen, M. Large language model-based agents for automated research reproducibility: An exploratory study in Alzheimer’s disease. arXiv:2505.23852 (2025). https://doi.org/10.48550/arXiv.2505.23852
- Michie, et al. Developing and using ontologies in behavioural science: addressing issues raised [version 2; peer review: 2 approved, 1 approved with reservations]. Wellcome Open Research 7, 222 (2023). https://doi.org/10.12688/wellcomeopenres.18211.2
- Elliott, H. et al. Living systematic reviews: An emerging opportunity to narrow the evidence-practice gap. PLoS Med. 11, e1001603 (2014). https://doi.org/10.1371/journal.pmed.1001603
- Borah, , Brown, A. W., Capers, P. L. & Kaiser, K. A. Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. BMJ Open 7, e012545 (2017). https://doi.org/10.1136/bmjopen–2016-012545
- Shojania, K. G. et al. How quickly do systematic reviews go out of date? A survival analysis. Intern. Med. 147, 224-233 (2007). https://doi.org/10.7326/0003-4819-147–4-200708210-00179
- Franco, , Malhotra, N. & Simonovits, G. Underreporting in psychology experiments: Evidence from a study registry. Social Psychological and Personality Science 7, 8-12 (2016). https://doi.org/10.1177/1948550615598377
- Laitin, D. et al. Reporting all results efficiently: A RARE proposal to open up the file drawer. Proc. Natl. Acad. Sci. U.S.A. 118, e2106178118 (2021). https://doi.org/10.1073/pnas.2106178118
- Butler, A. R., Hartmann-Boyce, J., Livingstone-Banks, J., Turner, T. & Lindson, N. Optimizing process and methods for a living systematic review: 30 search updates and three review updates later. Clin. Epidemiol. 166, 111231 (2024). https://doi.org/10.1016/j.jclinepi.2023.111231
- Kane, & Amin, B. Amending the literature through version control. Biol. Lett. 19, 20220463 (2023). https://doi.org/10.1098/rsbl.2022.0463
- Budd, M., Sievert, M., Schultz, T. R. & Scoville, C. Effects of article retraction on citation and practice in medicine. Bull. Med. Libr. Assoc. 87, 437-443 (1999).
- Suchak, T. et al. Explosion of formulaic research articles, including inappropriate study designs and false discoveries, based on the NHANES US national health PLOS Biol. 23, e3003152 (2025). https://doi.org/10.1371/journal.pbio.3003152
- Lin, FOCUS: An AI-assisted reading workflow for information overload. Nat. Biotechnol. 43, 2070-2075 (2025). https://doi.org/10.1038/s41587-025-02947-8
- Akhtar, et al. in Proceedings of the Eighth Workshop on Data Management for End-to-End Machine Learning. 1–6 (Association for Computing Machinery).
- Perkel, M. Make code accessible with these cloud services. Nature 575, 247-248 (2019). https://doi.org/10.1038/d41586-019-03366-x
- Ard, et al. Integrating data directly into publications with augmented reality and web-based technologies – Schol-AR. Scientific Data 9, 298 (2022). https://doi.org/10.1038/s41597-022-01426-y
- Puebla, et al. Ten simple rules for recognizing data and software contributions in hiring, promotion, and tenure. PLoS Comput. Biol. 20, e1012296 (2024). https://doi.org/10.1371/journal.pcbi.1012296
- Coalition for Advancing Research Agreement on reforming research assessment, <https://coara.eu/agreement/> (2022).
- Declaration on Research San Francisco Declaration on Research Assessment, <https://sfdora.org/read/> (2012).
- Brand, A., Allen, L., Altman, M., Hlava, M. & Scott, J. Beyond authorship: Attribution, contribution, collaboration, and credit. Learned Publishing 28, 151-155 (2015). https://doi.org/10.1087/20150211
Editors
Kathryn Zeiler
Alex Holcombe
Editorial assessment
by Alex Holcombe
Two of the original reviewers have looked at the revision and both are highly satisfied with this revision. The revision is substantial, having considered the reviewers’ comments extensively, and the manuscript is much improved. This is a significant contribution to the important and quite urgent topic of how scientific publication should change in response to the advances in AI.
Recommendations for enhanced transparency
- Add author ORCID iD.
- Add a competing interest statement. Authors should report all competing interests, including not only financial interests, but any role, relationship, or commitment of an author that presents an actual or perceived threat to the integrity or independence of the research presented in the article. If no competing interests exist, authors should explicitly state this.
- Add a funding source statement. Authors should report all funding in support of the research presented in the article. Grant reference numbers should be included. If no funding sources exist, explicitly state this in the article.
For more information on these recommendations, please refer to our author guidelines.
Peer review 1
Good job with the revisions
Peer review 2
Anonymous reviewer
I would like to thank the author for the detailed responses and for implementing the requested changes.
The revision fully addresses all the points raised in my initial review.
Consequently, I am satisfied with the current version.


