Market dynamics, governance and open research metadata in the AI era

Daniel W. Hook1 Email

1 Digital Science, 6 Briset Street, London, EC1M 5NR, UK

 

Originally published on April 24, 2026 at: 

Abstract

The debate about scholarly knowledge infrastructure has long been framed as a contest between openness and commercial enclosure. This framing distorts both policy and practice. The real tension lies between the persistent cost of producing and refining structured metadata under deep technological friction, and the differentiated demands distinct communities place on data quality, focus and granularity. We introduce the innovation annulus: the zone between freely available structured data and the advancing frontier of commercially refined knowledge products. This zone is a permanent, functional feature of the ecosystem -- not a pathology to eliminate. By analogy with the efficient market hypothesis, its width measures production inefficiency, set by the interplay of friction and demand. Artificial intelligence reshapes the annulus, lowering barriers to basic structuring, raising the threshold at which refinement adds value, and introducing systemic risks through unprovenanced AI-derived metadata. CRediT contributions, funding acknowledgements and AI disclosure statements illustrate the annulus lifecycle. Governance should calibrate the annulus, not abolish it: thin enough to serve research efficiently, wide enough to sustain innovation. A formal welfare framework, analogous to the Nordhaus optimal patent life, characterises the trade-offs and yields testable predictions. The Barcelona Declaration offers a promising forum for boundary governance.

I Introduction

For more than three decades, the debate about scholarly knowledge infrastructure has been organised around an opposition between openness and commercial enclosure. On one side, a sustained community of advocates, funders and policymakers has argued that the broadly defined products of publicly funded research including metadata about this research should be openly available [63]. On the other, commercial actors have maintained structured data products whose value depends, to varying degrees, on restricting access. The result has been a policy conversation framed as a zero-sum contest: Every advance in openness is a retreat for commercial interests, and vice versa.

This framing has been productive in some respects. It has driven the creation and development of open identifier infrastructure—Crossref, ORCID, ROR, DataCite—and the progressive opening of citation data and abstract data through initiatives such as I4OC and I4OA respectively. It has generated mandates for open access to published research and, most recently, the Barcelona Declaration on Open Research Information [9], which extends the logic of openness from publications to the broader research metadata ecosystem. These are significant achievements.

Yet this framing, at a deeper level, is now less productive—and its persistence has distorted both policy and practice [35, 13, 17]. We assert that the central tension in scholarly knowledge infrastructure is not between openness and commerce. It is between two features of the system that will not go away: First, the persistent and non-trivial cost of producing, structuring and refining knowledge data in a system characterised by deep technological frictions; and second, the demands that a variety of communities of users have for data to support a radically differentiated use cases that require vastly different data quality, focus and granularity.

These two features interact to produce what we call the innovation annulus [29]—a zone between the open core of structured data that is either free or indistinguishable from free for the user, and the advancing frontier of commercially refined knowledge products (see Fig. 1). As open access becomes more the dominant mode of publication, the annulus cannot be solely positioned as a pathology nor as an information asymmetry imposed by rent-seeking commercial actors. Rather, it is a structural consequence of production friction in a system where the scholarly record was never designed for structured data consumption, and where the definition of what constitutes useful structured data is continuously evolving. The width of the annulus is, in a precise sense, a measure of the system’s distance from perfect efficiency—and because perfect efficiency is logically impossible in a system where useful data types evolve faster than the production system can standardise them, the annulus is a permanent feature of the landscape. In the AI era in which we now sit, we might expect AI to reduce production friction, but given the copyright complexities of text and data-mining and the challenges of understanding the algorithmic provenance of a piece of data so that it can be trusted [57], it is not clear that this assumption is well-founded.

This paper develops this argument in several steps. We begin by examining the cost structure of scholarly data production and the technological frictions that sustain the annulus (Sec. II). We then analyse the differentiated demand that gives the annulus its sectoral dimensions, develop a geometric interpretation of the annulus diagram that yields diagnostic measures including an openness ratio for each data type (Sec. III), and present a formal welfare framework—analogous to the Nordhaus optimal patent life—for reasoning about the optimal annulus width (Sec. IV). We describe how artificial intelligence reshapes the annulus without eliminating it (Sec. V), and examine structured in-paper metadata—CRediT author contributions, funding acknowledgements, AI disclosure—as frontier data types that illustrate the annulus lifecycle (Sec. VI). We then turn to the governance question: not whether the annulus should exist, but how its boundaries should be managed (Sec. VII). We use the experience of Dimensions as an illustration of annulus dynamics in practice (Sec. VIII), and conclude with a set of open questions for the research community (Sec. IX).

II Production Friction and the Efficient Market Analogy

The production of structured scholarly data is often discussed as though it were a problem that technology has essentially solved, leaving only political and institutional barriers to full openness. This view underestimates the depth and persistence of the frictions involved.

To see why, it is helpful to borrow a concept from financial economics. The efficient market hypothesis (EMH) [19] holds that, in a perfectly efficient market, all available information is reflected in prices and no actor can gain systematic advantage from private information. The EMH is not a description of how markets actually behave; it is a benchmark—a theoretical limit against which the efficiency of real markets can be measured. Transaction costs, information asymmetries and regulatory frictions all cause real markets to deviate from the benchmark, and much of financial regulation is directed toward reducing these deviations.

We can apply the same logic to the scholarly data production system. In a perfectly efficient system—one in which all scholarly data were produced in a fully standardised, structured format at the point of creation, with universal identifiers, machine-readable metadata and complete provenance—there would be no need for downstream normalisation, enrichment or disambiguation. The information latent in the scholarly record would be fully expressed in its published form. The annulus would collapse to zero.

The actual scholarly data production system is, of course, very far from this benchmark. Scholarly data are produced by a highly heterogeneous publishing ecosystem that evolved over centuries without overall standardisation. Author names of individual researchers vary from publication to publication either through naturally arising inhomogenity or through more structured mechanisms such as the application of different publisher and journal house styles. Institutional affiliations are recorded inconsistently—a researcher may name their department but not their university, or use any of several variant forms (“University of Oxford” versus “Oxford University” versus “Dept. of Physics, Oxford”). Citation lists are formatted differently across thousands of journals. Funding acknowledgements may or may not include grant numbers, may name funders in full or use ambiguous abbreviations. These are not minor inconveniences; they represent deep structural frictions in the data production system that create a persistent demand for downstream normalisation [27].

Figure 1. The Innovation Annulus. The outer circle represents the structuring frontier—the frontier of what has been structured and made available in analysable form. Beyond the outer circle lies unprocessed, unstructured, or as-yet-unpublished scholarly output. The inner circle represents the boundary of openness boundary—data that is freely licensed, transparently governed, stable and available for unrestricted reuse. The zone between the two circles is the innovation annulus: Data that has been structured and refined to a level that makes it useful, but that has not yet been commoditised to the point of free availability. Arrows on the outer circle indicate expansion (new data types, new structuring capabilities). Arrows on the inner circle indicate outward migration (standards adoption, collective disclosure, technological commoditisation). The width of the annulus at any point represents the gap between the current frontier of commercially refined data and the current baseline of open provision. Adapted from Hook [29].

The annulus, by this account, is a direct measure of the system’s production inefficiency. Its width for any given data type reflects the gap between how that data type is actually produced (with all its frictions, inconsistencies and omissions) and how it would need to be produced for downstream use to require no additional investment. Where the gap is large—as it historically has been for author disambiguation, institutional identification and citation linking—the annulus is wide, and there is economic space for actors who invest in closing the gap. Where the gap is small—as it is becoming for basic bibliographic metadata decorated with DOIs and deposited through Crossref—the annulus narrows and the data migrates toward the open core.

Three categories of production friction sustain the annulus even as technology advances:

  1. Source heterogeneity. The scholarly record comprises journal articles, conference proceedings, books, book chapters, preprints, patents, clinical trial registrations, policy documents, datasets, software, peer reviews [65], seminars [66, 22], and more—each with different metadata conventions, identifier systems and quality standards. Structuring this heterogeneous landscape into a coherent, interlinked graph requires continuous investment in entity resolution, disambiguation and cross-referencing that scales with the diversity of the record rather than its volume alone [27, 55].

  2. Quality decay. Structured data does not remain accurate without maintenance. Researchers change institutions, journals change publishers, grants are reclassified, organisations merge or dissolve, and classification systems evolve. The cost of maintaining a living, accurate representation of the scholarly record is not a one-time investment but an ongoing operational expense—what Borgman [10] has characterised as the continuous labour of knowledge infrastructure stewardship, and what Edwards et al. [18] have documented as the invisible maintenance burden of knowledge infrastructures.

  3. Frontier data types. Even if every existing friction were resolved—if every author had an ORCID, every institution a ROR, every reference a DOI—the system would continue to generate new data types that begin their life in an unstructured, unstandardised state. Twenty years ago, nobody needed structured AI disclosure metadata because AI was not used at any significant scale in research. Ten years ago, few were thinking in terms of using structured research integrity signals at scale. The definition of what constitutes useful structured data is continuously expanding, and each new data type restarts the cycle of friction, investment, standardisation and eventual commoditisation.

This last point is critical. It means the annulus is not merely a legacy of historical infrastructure debt that will be paid down over time. It is a permanent structural feature of any knowledge system in which the definition of useful structured data evolves faster than the production system can standardise it. The annulus shrinks in specific places as standards mature—but it re-emerges at the frontier wherever new information needs outpace the system’s native production capacity.

A deeper physical grounding for this claim is available and worth making explicit. In The Human Use of Human Beings, Wiener framed organised systems—biological, mechanical, institutional, informational—as local enclaves that persist in defiance of the second law of thermodynamics only by continuously importing energy and exporting disorder to their surroundings [68]. Structured information is, in this view, a form of negentropy; any system that maintains it must continuously do work against the universal tendency toward disorder. Left to themselves, messages accumulate noise, classifications decay, identifiers decouple from the entities they name, and standards fragment as they are locally reinterpreted. The three categories of production friction identified above are specific expressions of this tendency rather than peculiarities of the scholarly ecosystem: source heterogeneity is entropy in the generative process; quality decay is entropy accumulating in a previously ordered store; and the continuous emergence of frontier data types reflects the expansion of the informational phase space faster than any one-time investment can contain. The annulus, from this perspective, is the thermodynamic signature of the work being done against that disorder. Its width for any given data type is, in effect, a measure of the energetic cost of holding that region of the scholarly record in an analysable state; eliminating it would require an impossible condition—an information system in which organisation required no ongoing input. Maxwell’s demon cannot get something for nothing, and neither can a metadata infrastructure: interventions that attempt to collapse the annulus do not remove the cost of maintaining order, they only reassign where that cost is paid. The annulus is the price of sustaining an anti-entropic enclave in an informational system that would, without continuous investment, drift toward unstructured noise.

The analogy with copyright is instructive here. Copyright creates a legally defined zone of exclusivity around creative works—a period during which the creator can recover their investment before the work enters the public domain. The width of this zone varies across domains (French moral rights for musical works extend further than Anglo-American economic copyright [52, 34]) and reflects a social and value-driven judgment about the importance of protecting creative effort. PhD thesis embargoes in arts, humanities and social sciences represent a similar mechanism: the community has collectively agreed that a graduating student’s interest in publishing a monograph justifies a temporary restriction on open access. In both cases, the width of the “annulus” is determined by a norm—legal, social or community-defined—that balances the public interest in access against the producer’s interest in return on effort.

In the case of scholarly metadata, no such norm exists. The width of the annulus is determined not by statute or community agreement but dynamically, by the interaction of production cost and market demand. This is not necessarily a problem—a dynamically determined width may be more responsive to real changes in production economics than any fixed norm—but it does mean that the width is vulnerable to market power. The historically inflated annulus of the Web of Science / Scopus duopoly era [42, 49] demonstrates what happens when the annulus width is sustained not by genuine production costs but by barriers to entry and institutional lock-in. The emergence of Dimensions [27], OpenAlex [58] and expanded open infrastructure through Crossref has demonstrated that much of what was in the annulus could be produced more cheaply and made available more openly—compressing the annulus toward a width that more accurately reflects genuine production friction.

It is worth noting that the annulus width for any given data type reflects not only technical production friction but also legal and contractual friction. Licensing restrictions on full-text access, for example, widen the annulus independently of the technical cost of mining. Even if AI were to reduce the computational cost of extracting structured metadata from full-text articles to near zero, the inability to access full text at scale—because of copyright restrictions, publisher licence terms, or paywalled content—keeps the outer boundary (the structuring frontier) further from the centre than the technical economics alone would dictate. A complete analysis of annulus width for any given data type must therefore decompose it into its technical component (the genuine cost of structuring) and its legal component (the artificial friction created by access restrictions). The governance prescriptions for reducing these two types of friction are quite different: technical friction is addressed by standards adoption and technological investment; legal friction is addressed by licensing reform, collective disclosure agreements and policy intervention.

III Differentiated Demand and the Sectoral Annulus

The annulus model as described so far treats the demand side as undifferentiated. In practice, the demand for structured scholarly data is radically heterogeneous, and this heterogeneity gives the annulus sectoral dimensions that move at different speeds and have different governance implications.

We identify three demand drivers that create distinct quality requirements, each of which shapes the width of the annulus in a different part of the knowledge spectrum.

III.1 Competitive strategy and information asymmetry

Research institutions compete for funding, talent and reputation. This competition is not a market distortion but a deliberate feature of research systems designed to allocate scarce resources to the most productive groups [16]. Competition creates demand for information asymmetries: institutions that can identify emerging research strengths more rapidly than and prior to their competitors, benchmark their performance more precisely, or anticipate funder priorities more accurately gain a structural advantage. This mechanism leads to the creation of centres of excellence and a varied landscape with specialisms. This dynamic is analogous to the knowledge spillovers that Jaffe [32] demonstrated in industrial R&D: structured data about who is doing what research, is the mechanism through which competitive intelligence flows, and actors who can access and interpret that data more effectively gain a systematic advantage.

This demand is not served by raw open data. It is served by refined, contextualised analytical products that transform the scholarly record into strategic intelligence. University research offices invest in benchmarking tools not because they object to openness but because their competitive environment rewards the ability to extract actionable insight faster and more accurately than other actors. The data they need is not merely accessible in the FAIR sense [69]—it must be Findable, Interoperable and Reusable in a specific institutional context, which currently requires curation and enrichment beyond what open provision typically delivers.

III.2 Research translation and alignment data

An increasingly important driver is the emerging ecosystem of research translation partnerships between academic and commercial organisations. The translation of basic research into application—what Stokes [62] mapped as “use-inspired basic research”—requires both parties to understand the alignment of their respective capabilities. In what Adams [3] has characterised as the “fourth age of research,” where the dominant mode of production is international collaboration, the data infrastructure required to support translation must be correspondingly global and interlinked. Which research groups are working on problems relevant to a company’s R&D pipeline? Which institutions have the expertise and infrastructure to support a translational challenge? These questions require high-quality data on research capabilities that goes beyond publication counts: entity-resolved researcher profiles, institutionally contextualised output data, patent-to-publication linkage, clinical trial mapping and funding flow analysis—all at a quality sufficient to support investment decisions [55].

III.3 Corporate research and domain-specific refinement

Research-intensive industries—biomedical and pharmaceutical most prominently, but also automotive, engineering, agrichemical, petrochemical and financial, to name just a few—conduct substantial research programmes and depend on the scholarly record for competitive intelligence, prior art analysis and strategic planning. Their data needs differ from those of academic users in important respects: they require data at a granularity and domain specificity that reflects their particular R&D portfolios; they are willing to pay for quality because the decisions that depend on it have direct financial consequences; and they often require integration across scholarly and non-scholarly sources—patents, regulatory filings, clinical trial registrations, market data—that falls outside the scope of what open scholarly infrastructure was designed to provide. Even at the national policy level, analyses demonstrating that research impact is driven primarily by international collaboration rather than domestic performance [2] require data structured and refined to a quality that supports robust interpretation [64]—a quality level that raw open metadata does not yet consistently deliver.

III.4 Reading the annulus geometry

The sectoral annulus model introduced above can be made more precise by attending carefully to the geometry of the diagram (Fig. 2). Each data-type segment has two independent visual properties: its radial position (how far from the centre both the inner and outer arcs sit) and its thickness (the gap between the inner and outer arcs). Together, these encode a surprisingly rich set of inferences about the state of a data type in the knowledge ecosystem.

Figure 2. Sectoral Dimensions of the Annulus. The concentric-circle diagram from Figure 1, but with the annulus divided into radial sectors of different widths and different radial positions, each labelled with a data type. Critically, the inner and outer arc of each sector are positioned independently. The radial distance from the centre represents the total structuring investment and ecosystem development for that data type: segments further from the centre indicate more mature, more developed, and more extensively structured domains. The thickness of each segment represents the commercial opportunity—the gap between what is openly available and what has been structured at the frontier.

The radial position of a segment captures the total structuring investment and ecosystem development for that data type. If both arcs are close to the centre, the data type is either embryonic (nobody has invested much in structuring it because demand has not yet materialised) or inherently limited in scope (the information content is bounded and does not require deep development). If both arcs are far from the centre, the data type has a well-developed ecosystem with significant structuring effort and mature systems. Radial position is not, however, purely about demand: a data type may be close to the centre either because there is little interest in it or because the cost of mining it is prohibitively high despite significant latent demand. The distinction matters for governance: the first case requires no intervention, while the second may require public investment or standards development to unlock the latent value.

The thickness of a segment represents the commercial opportunity—the gap between what is openly available and what has been structured at the frontier. A thick segment indicates a large space in which commercial refinement can operate. A thin segment indicates that open provision is close to the frontier, leaving little room for commercial differentiation.

These two properties interact to create four characteristic configurations:

  1. A segment that is thin and close to the centre represents a data type that is low-cost to produce, for which adequate systems are in place, and where little innovation is required. Structured peer review metadata might currently fall into this category—some data exists, but neither demand nor production investment is deep.

  2. A segment that is thin and far from the centre represents the success case: mature systems, good standards, and open provision that has caught up with the frontier. Basic bibliographic metadata with DOIs is approaching this state. The annulus has been compressed by standards adoption and collective disclosure through Crossref.

  3. A segment that is thick and close to the centre represents either gatekeeping (significant structuring has occurred but little is openly available) or an embryonic field where the data are expensive and difficult to mine. The legacy Web of Science model exemplified a gatekeeping pattern: the inner arc was very close to the centre (almost nothing open) while the outer arc was moderately distant (substantial commercial product). This makes sense during the initial development of a new data source such as the Web of Science. In the 1950s the data were not diverse, it was hard to mine and hence a homogeneous annulus would be an appropriate representation similar to Figure 1). The openness boundary can be shrunk to the centre with the annulus becoming a fully filled circle as there was essentially no open data—the cost of production precluded that option at that time—the data were excessively expensive to mine. Indeed, so expensive was it to mine data that it was only Garfield’s critical realisation was that citations were Pareto distributed [11, 21] that made it financially tractable to construct an index [47, 71]. For more recent emerging data types like AI disclosure, modern technologies automatically set the boundaries differently—the thickness reflects tends to be closer to the market expectation of the production cost.

  4. A segment that is thick and far from the centre indicates a data type with extensive development at the frontier and a substantial commercial zone, but also a meaningful open core. Domain-specific enrichment—pharmaceutical patent-to-publication linkage, translational alignment data, institutional analytics—sits here. The thickness reflects the genuine cost of refinement at the quality levels that demanding users require.

Several additional inferences follow from this geometry. First, the ratio of the inner radius to the outer radius for each segment yields a natural openness ratio—the proportion of total structuring effort that is openly available. A ratio approaching one indicates near-complete openness for a mature data type. A ratio close to zero indicates that most structured data remains behind a commercial boundary. This ratio is a useful diagnostic for governance: one might argue that for data types essential to basic research management, the openness ratio should be above some threshold, even if the absolute frontier extends further for specialised users.

Second, the diagram should be understood as a snapshot of a dynamic system. The trajectory of each segment is as important as its current position. A segment in which both arcs are moving outward indicates a dynamic, maturing data type—new structuring is occurring and new data is becoming open. A segment in which the outer arc moves outward while the inner arc remains stationary represents increasing commercial development without corresponding openness gains—a pattern that should raise governance concerns. A segment in which the inner arc is catching up with the outer represents active commoditisation. And a segment in which neither arc moves is stagnant.

Third, there is a structural constraint that the diagram makes visible: the inner boundary (or openness boundary) can never move outward faster than the outer boundary (the structuring frontier) in the long run, because data cannot be made openly available until it has been structured. Open provision is bounded above by total structuring effort. This has a crucial implication for the argument of this paper. In domains where no one is investing in frontier structuring—because there is no commercial incentive and no public funding—the inner boundary stalls too. The annulus does not merely describe the commercial gap; it describes the investment incentive. An annulus that is too thin removes the economic incentive for frontier structuring, which in turn slows the outward movement of the inner boundary. This is the strongest version of the argument that the annulus is functional: without sufficient annulus width, the entire system—including the open core—advances more slowly.

Fourth, there is an important category of cases in which the inner arc is further out than market forces alone would produce. This represents deliberate intervention—a mandate, a philanthropic investment, or a strategic commercial decision to make data open beyond what the natural economics would deliver. The Dimensions free tier and BigQuery access [24, 28] represent strategic commercial decisions to push the inner boundary outward. OpenAlex [58] represents philanthropic funding achieving a similar effect. The Barcelona Declaration [9] is, in effect, an attempt to create coordinated pressure to push the inner boundary outward across many data types simultaneously. The inner boundary is not solely determined by production cost—it is also shaped by values, strategy and governance choices.

III.5 The equity dimension

A consideration that cuts across the three demand drivers and the geometric analysis above concerns the equity implications of how the annulus is structured. In a world where AI can be used to enhance and enrich base-level data, institutions with computational resources, technical talent and data science capacity can run their own enhancement pipelines—resolving affiliations, classifying output, identifying collaboration opportunities. Institutions in under-resourced settings cannot. The data is nominally open, but the capacity to extract value from it is profoundly unequal.

This creates a strong case for centralised provision at accessible cost. If each institution must independently enhance its own data, the gap between rich and poor institutions widens. If intermediaries—whether commercial or community [24]—produce enhanced, trustworthy data that everyone can agree is a source of “truth” for certain evaluative activities at a consistent quality standard, the baseline rises for everyone. The centralised model is not just more computationally efficient; it is more equitable, because it sets a floor below which no institution needs to fall [15]. Lane [40] has articulated this principle forcefully in the context of public data systems: Data infrastructure should be designed so that access does not depend on insider knowledge or privileged networks, because “unless a researcher is able to tap into a network of cognoscenti, they would not know about the data, or not know how to use it” [39]. The same logic applies to scholarly metadata: the open core must be not merely technically accessible but practically usable by institutions at all resource levels. Adams, Gurney, Hook and Leydesdorff [1] have demonstrated what structured data can reveal about collaboration patterns in Africa—analyses that would be impossible without the kind of institutionally resolved, openly available metadata that the inner circle of the annulus is designed to provide. More recently Pinfield [54] argues that nominal openness isn’t sufficient when the capacity to extract value from open data is unequally distributed.

The governance implication is that public investment should focus on ensuring that the baseline of structured data available to all institutions is as high as possible—not on eliminating the annulus entirely. The annulus above the baseline serves differentiated demand from users who can afford to pay for frontier refinement. The baseline—the inner circle—serves the equity function of ensuring that under-resourced institutions are not excluded from the structured data they need for competent research management and strategic decision-making. In some sense, this constitutes a modern form of institution building—in this case not research-performing institutions but rather social institutions that define norms and expectations of data, its quality and its uses.

IV Toward an Optimal Annulus Width

The geometric framework developed in Sec. III.4 makes the annulus visually tractable. The natural next question is whether it can be made analytically tractable: Is there a principled way to determine what the annulus width should be for a given data type, rather than merely observing what it is?

The question has a structural analogue in the economics of intellectual property. Nordhaus [50] posed the same question for patents: given that a period of monopoly exclusivity is needed to incentivise invention, what is the optimal length of that period? His answer—a formal optimisation trading off the deadweight loss of monopoly pricing against the incentive to invest in R&D—established a framework that has been extended by Klemperer [37], Gallini [20] and Scotchmer [59], and that remains foundational in innovation economics. We argue that the annulus presents a structurally homologous problem, and that a similar framework can be developed for scholarly knowledge infrastructure.

IV.1 A welfare framework

Consider a single segment of the annulus at a given point in time. Let ri denote the inner radius (the openness boundary) and ro the outer radius (the structuring frontier). The annulus width is w=ro−ri and the openness ratio is ρ=ri/ro.

It is helpful to define the value and cost functions in an order that parallels how data enter the system: structuring first, and open provision as an overlay on structured data. We therefore begin with the total value and cost of structuring, and then introduce separately the additional welfare and different cost profile associated with making part of the structured data openly available.

Value of structuring. Let V​(ro) denote the total value, to all users, of having data structured up to frontier level ro, evaluated at whatever access terms apply (paywalled, tiered or free). V is increasing and concave in ro: deeper structuring yields diminishing returns as progressively more specialised data types are brought into the structured domain.

Cost of structuring. Let C​(ro) denote the full cost of producing and sustaining structure up to level ro—entity resolution, disambiguation, classification, and the ongoing stewardship required to keep the structured representation current as the underlying record evolves [18, 10]. C is increasing and convex: the easy structuring tasks (for example, applying DOIs, resolving common institutional names) are accomplished first, and each additional unit of frontier structuring requires more domain expertise, more validation and continuing maintenance. We do not separate production from maintenance because they are incurred jointly and scale together with ro.

Openness premium. Let B​(ri) denote the additional welfare gain from making the range [0,ri] openly available rather than providing it on restricted terms—the equity, access and knowledge-spillover premium that openness is incremental to the value captured in V. B is increasing and concave in ri: the first units of open provision (basic bibliographic metadata, core identifiers) have enormous marginal value because they set a floor for all institutions, while later units have diminishing marginal impact. This is where the equity argument of Sec. III.5 enters formally—the social weight on the early units of B is high because they serve under-resourced institutions that would otherwise be excluded [61].

Cost of the openness overlay. Let M​(ri) denote the additional cost of running the openness overlay on the range [0,ri]: data standardisation and governance, free distribution at scale, community oversight, and the persistence and preservation guarantees that openness requires. M is increasing in ri and is subject to external pressures that may raise it exogenously—the AI-harvesting cost inflation documented by Crossref [23] enters through M rather than C.

Under this decomposition, paywalled data inside the annulus incurs C but not M; open-core data incurs both. The social welfare function is then

W = V​(ro) + B​(ri) − C​(ro) − M​(ri).         (1)

This is subject to a sustainability constraint: the system must be financially viable. In general, revenue depends on both boundaries, since deeper frontier refinement commands higher marginal revenue per unit of width than shallow refinement of commoditised data. The appropriate object is therefore R​(ri,ro). Within a single segment of the annulus, where the radial range is narrow, revenue is locally well approximated as a function of width alone, and we write R​(w) with the understanding that this is a local reduction. We return to this point in Sec. IV.3. Revenue plus any public subsidy S must cover the costs:

R​(w) + S ≥ C​(ro) + M​(ri).         (2)

The constrained optimisation yields first-order conditions that characterise the optimal boundaries. For the inner boundary,

B′​(ri) = (1 + λ) ​[M′​(ri) + R′​(w)],         (3)

and for the outer boundary,

V′​(ro) = (1 + λ) ​[C′​(ro) − R′​(w)],         (4)

where λ≥0 is the shadow price of the sustainability constraint and a dash implies the derivative of the function with respect to its natural variable, thus B′​(ri)=d​B/d​ri. λ measures the social welfare gained from relaxing the constraint by one unit—for example, through an incremental unit of public subsidy. When revenue plus subsidy more than covers costs, the constraint does not bind and λ=0; the first-order conditions reduce to the unconstrained conditions B′​(ri)=M′​(ri)+R′​(w) and V′​(ro)=C′​(ro)−R′​(w). When the constraint binds tightly, λ is large and the annulus must do more of the work of sustaining the system.

The first-order conditions have a direct intuitive reading. Equation (3) states that the marginal social value of expanding the open core, B′​(ri), should equal the marginal cost of doing so—directly, through the incremental cost of open provision, M′​(ri), and indirectly, through the revenue forgone by narrowing the annulus, R′​(w). The weight (1+λ) reflects the fact that, when money is tight, costs and forgone revenues count more heavily against social welfare. Equation (4) is the symmetric condition for the outer boundary: the marginal value of frontier structuring, V′​(ro), should equal the marginal structuring cost, C′​(ro), less the marginal revenue gained from widening the annulus, R′​(w), again weighted by the tightness of the financial constraint. Readers seeking a textbook treatment of the underlying constrained-welfare machinery will find the most relevant material in Stiglitz [61] and, at greater technical depth, in Atkinson and Stiglitz [8].

Figure 3 explores these conditions pictorially. Panel (a) plots net welfare W as a function of annulus width w, holding the outer boundary at its optimal level: the curve is an inverted-U with maximum at w∗. As w→0 the sustainability constraint binds and frontier investment becomes unsustainable; as w grows large the social benefit of openness is progressively forgone. The asymmetry of the curve is deliberate: the left shoulder reflects a discrete feasibility collapse (the constraint ceases to hold), while the right shoulder reflects a continuous opportunity cost (forgone units of B​(ri)). Panel (b) shows the boundaries ri and ro and the width w of the annulus.

Figure 3. A welfare-theoretic view of the optimal annulus width. Panel (a) plots net welfare W as a function of annulus width w, holding the outer boundary at its optimal level. The inverted-U is asymmetric: the left shoulder reflects a discrete feasibility collapse as the sustainability constraint, Eq. (2), ceases to hold and frontier structuring becomes unsustainable; the right shoulder reflects the continuous opportunity cost of forgoing units of B​(ri). Panel (b) shows the boundaries ri and ro and the width w.

IV.2 Qualitative predictions

Even without solving the optimisation for specific functional forms, the framework yields several qualitative predictions that are, in principle, testable.

Prediction 1: the optimal annulus width is wider where structuring costs are higher. When the marginal cost (i.e. the cost of obtaining an additional data point in a given set) of structuring data C′​(ro) is large, the sustainability constraint binds more tightly (higher λ), and more revenue from the annulus is needed to sustain frontier investment. This explains why the annulus persists for data types that require deep domain-specific enrichment (pharmaceutical patent linkage, translational alignment data) while narrowing for data types where AI has reduced structuring costs.

Prediction 2: the optimal width is narrower where the social benefit of openness is steep. When the marginal welfare gain (i.e. the welfare associated with making open an additional data point in a given data set) B′​(ri) is large—for data types essential to basic research management, where equity concerns dominate—the first-order condition (3) pushes ri outward, narrowing the annulus. This formalises the intuition that governance should prioritise high openness ratios for foundational data types.

Prediction 3: the optimal width shrinks as technology reduces structuring costs. As AI and standards adoption reduce C​(ro), the sustainability constraint loosens (lower λ), and the inner boundary can move outward. This is the formal version of the claim that technology compresses the annulus—but it also predicts that the compression is modulated by the shadow price, not automatic. There are a variety of assumptions implicit in this. Two of the most significant are that: i) the externalities associated with data labelling and the environmental impact of AI are negligible; ii) that algorithmic data improvement finds a systematic structure that allows context of data improvements to travel with metadata in a transparent manner.

Prediction 4: public subsidy is a substitute for annulus width. A larger S relaxes the sustainability constraint, allowing a thinner annulus at the same level of frontier investment. The Entrepreneurial State model [43] corresponds to the limiting case S→S∗ where subsidy fully replaces annulus revenue and w→0. The functional annulus model corresponds to moderate S and moderate w. The framework makes explicit what each model asks of the public purse. The risk of this model is that innovation may lack an appropriate efficiency moderator.

Prediction 5: an annulus that is too thin slows the entire system. This follows from the structural constraint that the inner boundary cannot advance faster than the outer boundary: ri≤ro at all times. If the annulus is compressed below the level at which frontier investment is sustainable (R​(w)+S<C​(ro)+M​(ri)), the outer boundary stalls—and with it, the inner boundary. Open provision is bounded above by frontier structuring. This is the formal expression of the argument that the annulus is functional: eliminating it does not merely reduce commercial revenue but reduces the rate at which the open core can expand.

IV.3 Relation to patent theory and limitations

The framework is structurally homologous to the Nordhaus optimal patent life derivation [50], but differs in two important respects. First, in the patent case, the deadweight loss arises from monopoly pricing during the exclusivity period—a dynamic that Jaffe and Lerner [31] have documented can lead to patent systems in which exclusivity periods bear little relationship to the investment required. In the annulus case, the analogous loss arises from restricted data access: institutions that cannot afford refined data make worse decisions than they would with full access, and the research ecosystem as a whole operates below its potential [7]. Second, patent life is uniform across all inventions within a jurisdiction. The optimal annulus width is a function of the data type’s characteristics—its structuring cost profile, its demand heterogeneity, and the social weight placed on open access to it. The framework does not yield a single optimal width but an optimal width function w∗​(C′,B′,V′,λ) that varies across segments of the annulus diagram. This is more complex than the patent case but also more realistic, since different data types in the scholarly ecosystem clearly operate under different economic conditions.

The framework has important limitations. First, the functional forms of B, V, C, M and R are not known empirically for any data type in the scholarly ecosystem; their estimation is itself a research programme (see Sec. IX). Second, the shadow price λ is not directly observable; proxies would need to be developed—for example, the ratio of unmet demand for open data to available public funding could serve as a coarse indicator of when the constraint binds tightly. Third, the reduction from R​(ri,ro) to R​(w) is a local approximation, valid within a single segment of the annulus where the radial range is narrow. Across segments, and at very different radial positions, revenue depends on the absolute level of structuring as well as on its differential from the open core, and the general form R​(ri,ro) is appropriate; the qualitative predictions of Sec. IV.2 are robust to this generalisation because they depend on signs of first and second derivatives rather than on the specific reduction to width. Fourth, the model is static: it characterises the optimal annulus at a point in time rather than the optimal trajectory of both boundaries. A dynamic extension—casting the problem as an optimal control problem with equations of motion for ri​(t) and ro​(t), driven by technological change and standards adoption—would be a natural next step, and would connect to the broader literature on the optimal timing of technology diffusion.

We note these limitations not to undermine the framework but to identify the empirical and theoretical work required to make it operational. Even in its current form, the framework provides something that the governance conversation around open research information has lacked: a formal language for reasoning about the trade-offs involved in setting the boundary, and a set of testable predictions about how the annulus should respond to changes in technology, demand and policy.

V AI and the Annulus

Artificial intelligence has transformed the economics of the annulus, but the nature of the transformation is frequently mischaracterised. The common narrative holds that AI will collapse the annulus entirely by automating metadata extraction, entity resolution and classification to near-zero cost. The framework developed in Sec. IV shows why this is too simple: AI reduces C​(ro), which loosens the sustainability constraint and allows the inner boundary to move outward, but it does not drive C to zero and it simultaneously affects M, V and the demand structure. Three specific dynamics warrant attention.

V.1 Frontier acceleration and cost reduction

AI techniques—automated metadata extraction, entity resolution, semantic classification, citation parsing—have lowered the cost and increased the speed of basic data structuring. Tasks that previously required teams of manual curators can now be accomplished computationally at a fraction of the cost. This is unambiguously positive: it pushes the inner boundary of the annulus outward, expanding the volume of data that can be provided openly, with “appropriate provenance”.

Additionally, the distribution of benefit is uneven. AI is most powerful when applied to large, well-organised data collections. Actors who already hold substantial structured data assets gain disproportionate advantage because their existing assets provide training data, ground truth and computational context [60].

V.2 Quality threshold elevation

A less-discussed effect of AI is that it raises the quality threshold at which data refinement has commercial value. As basic structuring becomes commoditised through AI automation, the annulus does not disappear—it migrates to higher-order refinement tasks. The commercially valuable frontier moves from “can you structure this data at all?” to “can you structure it at a quality level that supports investment decisions, regulatory submissions or competitive strategy?”

This dynamic is familiar from other technology-intensive industries. In financial data, the commoditisation of basic market data did not eliminate the market for refined analytics; it created a new premium tier of algorithmic intelligence. In geospatial data, the availability of free satellite imagery did not eliminate demand for domain-specific analysis; it shifted the value frontier from data acquisition to data interpretation. The same pattern is visible in scholarly data.

V.3 Systemic risks of unprovenanced AI-derived metadata

AI also introduces a systemic quality risk that has received insufficient attention. When multiple actors independently use AI to enhance metadata without tracking the provenance and processing history of their enhancements, the result is not merely duplicated effort—it is a potential quality degradation. An AI-derived institutional affiliation that is contributed back into a shared system without its processing history becomes, for the next consumer, an apparently authoritative data point whose reliability cannot be assessed. If that consumer’s AI then builds on it, errors compound. Porter [56, 57] has articulated this risk clearly: AI enhancement of metadata without understanding or context can lead to poorer quality data, as downstream users do not understand the full processing and provenance of a piece of metadata that has been contributed to a centralised system without the details of its history.

This creates a system-level efficiency argument for centralised structured data intermediaries—whether open or commercial—that goes beyond the usual access debate. The intermediary exists not to restrict access but because centralised normalisation with provenance tracking is more efficient and more reliable than distributed rederivation without it. In an era of concern about computational carbon footprints, the duplication cost of many independent AIs repeatedly rederiving the same structured data—rather than consuming it from a maintained, provenanced central store—is itself a consideration.

VI Structured In-Paper Metadata: A Frontier Case Study

To illustrate the annulus lifecycle concretely, we examine a set of data types that are currently at different points on the journey from unstructured free text to standardised, openly available metadata. These are not data about papers derived by external analysis (retraction databases, image manipulation detection) but data produced by authors as part of the publication process that are not yet systematically structured into the formal scholarly record.

Funding acknowledgements. Information the funding of a piece of research is present in the vast majority of papers but is structured in wildly inconsistent ways. Some authors include full funder names and grant numbers; others use ambiguous abbreviations; still others include only a brief narrative acknowledgement. Crossref’s funder registry has begun to standardise this, and publishers increasingly request structured funding information—but the legacy record and the inconsistency of current practice mean that producing reliable, analysis-ready funding data at scale requires significant investment in natural language processing and entity resolution.

CRediT author contribution statements. The CRediT taxonomy [6, 12, 46] provides a standardised vocabulary for describing author contributions (conceptualisation, methodology, writing, supervision, etc.). CRediT was designed to address the inadequacy of ordered author lists as a mechanism for attribution and credit [5]. Adoption is growing but uneven. A recent retrospective analysis found that as of 2024, only 22.5% of original research articles with available full text in Dimensions included CRediT role information, with significant variation across publishers, disciplines and countries [4]. Where CRediT statements are present, they are not always machine-readable; where they are machine-readable, they are not always deposited in Crossref metadata. The Dimensions team, for example, has invested in the creation and curation of AI models that identify author contribution statements across the literature—work that operates at accuracy levels that still require improvement and hence further investment [29]. These data would be of significant value to the evaluation community and to anyone involved in tenure and promotion processes, but no universally accepted structured data format yet makes them widely available.

Data availability statements. Many journals now require authors to declare whether and where their research data are available. These statements are typically free-text, with no standard vocabulary or structure. Extracting structured information about data availability at scale, particularly distinguishing between “data available on request,” “data deposited in [specific repository],” and “no data were generated” requires significant processing.

AI disclosure. The most recent addition to this category, AI usage disclosure is currently required by a growing number of journals but in entirely unstandardised forms. The disclosure may appear in the methods section, the acknowledgements, a dedicated statement, or nowhere at all. There is no agreed vocabulary for describing what role AI played (writing assistance, data analysis, code generation, image creation). This represents a data type at the very outer edge of the annulus: enormously valuable for understanding the evolution of research practice, but currently extractable only through expensive, error-prone natural language processing.

Each of these data types follows a recognisable lifecycle. They begin as unstructured free text within papers (wide annulus, high production cost for anyone wanting to use them at scale). They move through a standardisation phase in which a taxonomy or identifier system is developed (CRediT, Crossref funder registry). They pass through a phase of uneven adoption in which the standard exists but is not universally applied. And eventually—though this has not yet happened for most of these data types—they become part of the baseline structured record that publishers produce natively, when the annulus for that data type shrinks toward zero.

The progression is real but slow, and it is driven by the same forces that the EMH analogy identifies: standards adoption that reduces production friction, collective disclosure agreements that expand the open baseline, and technological advancement that lowers the cost of extraction and normalisation. Cross-publisher efforts to standardise research data policies [30] and emerging initiatives to rethink publication models around open science principles [36] represent governance interventions that accelerate this progression by reducing production friction at source. Research linking publications to deposited data has demonstrated measurable citation advantages [14], providing empirical evidence that structured metadata creates value—and hence that the investment in structuring is justified. It is worth noting that if originators—authors and publishers—had perfectly structured these data at source, with standard taxonomies and machine-readable formats, the normalisation annulus for these data types would not exist. The annulus is, in this precise sense, a consequence of historical and ongoing production inefficiency. As the publishing system modernises, and unique identifiers and structured data standards become more universal, the cost of data production and maintenance for these established data types will decrease. But new data types that the ecosystem will then want to consume—signals that we cannot yet anticipate—will create new frontiers where the same dynamic plays out again.

VII Governance: Managing the Boundary

If the annulus is a permanent structural feature of the knowledge ecosystem rather than a pathology to be eliminated, then the governance question changes. It is no longer “how do we make everything open?” but “how do we ensure the annulus is the right width?” Too thin, and there is insufficient economic incentive for the frontier data production that serves differentiated demand and drives innovation in data quality. Too thick, and the research ecosystem is poorly served—locked into paying for data above its actual value.

VII.1 The Entrepreneurial State and its limits

The most developed case for public intervention in knowledge infrastructure is Mazzucato’s Entrepreneurial State [43], extended in her work on mission-oriented policy [44] and value theory [45]. Mazzucato argues that the state is not merely a passive corrector of market failures but an active co-creator of markets and technologies, and that publicly funded research constitutes a public investment whose returns should accrue to the public. Open knowledge infrastructure fits naturally into this framework: if public funding produces the research, the state should invest in the infrastructure required to make that research findable, usable and assessable as a public good. This argument has significant force. State investment has produced genuine public goods—from identifier infrastructure (ORCID, ROR) to national open science initiatives—and the Mazzucato framework provides the strongest theoretical justification for continued public investment in the open core of the annulus.

Moreover, the question of how to track and measure the returns on public investment in research infrastructure—which Lane, Owen-Smith and Weinberg [41] have recently addressed in the context of AI—is itself dependent on the kind of structured, linked metadata that the annulus model describes. The Entrepreneurial State cannot assess whether its investments are generating public value without the data infrastructure to measure outcomes, creating a recursive dependency: the state needs structured data to justify its investment in structured data.

But applying the Entrepreneurial State Model (ESM) to the full spectrum of data refinement needs creates difficulties that the annulus framework makes visible. The welfare framework in Sec. IV makes the trade-off explicit: the ESM corresponds to the limiting case in which public subsidy S fully replaces annulus revenue and the annulus width w approaches zero. This is logically coherent but practically demanding for three reasons.

First, cost allocation. In an idealised scenario, public institutions could centralise the refining process so as to avoid inefficiency and duplication of effort. However, in bearing the full cost of refining data to the quality levels required by pharmaceutical companies, venture capital firms and corporate R&D departments, the public subsidy flows disproportionately to private beneficiaries. This is not market creation—the canonical justification for the Entrepreneurial State—but public subsidy of private competitive intelligence [48]. The distinction matters: Mazzucato’s argument is strongest when the state creates infrastructure that enables private innovation (roads, internet protocols, basic research); it is weaker when the state produces the specific refined products that private actors would otherwise pay for.

Second, fiscal fragility. State funding is subject to political cycles and fiscal pressures, creating structural vulnerability. A knowledge infrastructure entirely dependent on public funding is exposed to exactly the kinds of budgetary shocks that the Mazzucato framework seeks to prevent in other domains. The history of state-funded data infrastructure is not reassuring: databases have been defunded, privatised, or allowed to decay when political attention shifts. The annulus model suggests a more resilient arrangement in which the open core is sustained by a combination of public investment and revenue from commercial frontier activity, diversifying the funding base rather than concentrating it in a single source.

Third, geopolitical risk. In a multipolar geopolitical environment, state-funded knowledge infrastructure carries risks of what Jasanoff [33] has termed divergent “civic epistemologies”—different political cultures’ ways of establishing what counts as reliable knowledge. A knowledge commons anchored to any particular state, or even a coalition of like-minded states, risks encoding particular epistemological assumptions into what presents itself as a universal infrastructure [15]. The concentration of infrastructure investment in the Global North has already produced a scholarly record with significant geographic biases; extending state-funded provision without addressing these structural biases may intensify rather than resolve the problem.

The annulus framework suggests a middle path: the Entrepreneurial State model is the right approach for the open core (foundational metadata, identifier infrastructure, the baseline that all institutions need), while the annulus provides the economic space for frontier refinement that serves differentiated demand and whose costs should not be socialised across the public purse. The governance challenge is ensuring that the boundary between these two zones is set in the public interest rather than by market power alone.

VII.2 The Crossref mechanism

An alternative governance mechanism—one that has received less theoretical attention than it deserves—is already operating in part of the scholarly data ecosystem. Crossref, orignally as an industry collaboration among publishers and now expanding to be more inclusive in its membership, functions as a de facto boundary-setting institution for one zone of the annulus. When Crossref expands the scope of standard metadata disclosure—requiring, for example, that deposited records include reference lists, abstracts or ORCID identifiers—it effectively moves a data type from the annulus into the open baseline. This does not happen through state mandate or market competition but through collective agreement among the publishers of the underlying data.

This is a genuinely distinctive institutional mechanism. It is not state provision, not market dynamics, and not community self-organisation in the Ostrom [51] sense. It is collective disclosure by an organised industry body, and it has been remarkably effective at expanding the open core for data types that publishers already produce. Crossref has an increasingly diverse membership including publishers, research institutions, funders and governmental organisations. But the observation that Crossref has, without anyone having designed it for this purpose, become a boundary-setting mechanism for one part of the annulus raises an important question: under what conditions might similar mechanisms emerge for other data types? And might Crossref itself, with an appropriately expanded governance structure, play a broader role?

VII.3 The Barcelona Declaration as a norm-setting forum

The Barcelona Declaration [9, 38] introduces a governance logic that is compatible with the functional annulus model. Rather than mandating specific data to be free, the Declaration frames openness as a civic obligation—what Porter [56, 57] has called “research information citizenship”—and distributes responsibility across producers, consumers and aggregators of metadata. Waltman [67] has argued that responsible research assessment requires open scholarly metadata—a position that the annulus framework refines: the question is not whether metadata should be open in principle, but which metadata types should be inside the open core at any given stage of the system’s maturation.

This framing is significant because it positions the Declaration not as a regulation but as a norm-setting body. In the same way that the scholarly community has developed informal norms about PhD embargoes and data sharing, the Barcelona Declaration could be the venue where the community develops norms about which data types should be inside the open core, what quality standards apply to open metadata, and what responsibilities different actors have in the production chain. The Declaration’s membership includes funders, institutions, and infrastructure providers—arguably the right constituency to develop these norms, though one could ask whether it should expand to include the corporate research users whose data needs shape part of the annulus.

Assessed against Ostrom’s [51] design principles for commons governance, the Barcelona Declaration framework has both strengths and gaps. It establishes boundaries and articulates proportional responsibilities. But it lacks effective monitoring of compliance, graduated sanctions for non-compliance, and formal conflict resolution mechanisms. The history of analogous declarations—the Budapest Open Access Initiative, DORA, Plan S—suggests that declarations without enforcement often fail to change institutional behaviour at scale [25, 70]. However, it is not necessary for there to be explicit, direct enforcement associated with these initiatives for them to be of value. Rather it is the engagement created by these approaches that engender change and new consensus in the policy environment. These initiatives do have each led to policy changes at institutional, local and national levels. Thus, ensuring that they provide inclusive mechanisms to host debate and strength their arguments to make them more potent in the policy arena would seem to be an important facet of their function.

VII.4 Open access as a parallel case

The open access experience provides a brief but instructive parallel. Open access mandates attempted to collapse the annulus in scholarly publishing by requiring that publicly funded research be freely accessible. What followed was not the elimination of commercial logic but its displacement: article processing charges preserved publisher revenues while shifting their form, and in some analyses total costs to the research community increased, with the burden redistributed regressively [53, 35, 17, 13]. The annulus did not disappear; it migrated from access charges to publication charges while its structural function remained intact.

The annulus framework predicts this outcome: as long as there are genuine costs in the publication production system that exceed what can be funded through public subsidy alone, an annulus of some form will persist. The lesson for research metadata governance is that policy which treats the annulus as a pathology produces displacement rather than resolution. The functional annulus model avoids this trap by accepting the structural reality and focusing governance energy on the boundary conditions rather than on the elimination of commercial activity within the annulus.

VIII Dimensions: An Illustration of Annulus Dynamics

Dimensions, the bibliographic database operated by Digital Science, provides an empirical illustration of annulus dynamics in practice. We describe it here not as a model to be copied but as a case study that makes several of the paper’s theoretical claims concrete.

Dimensions was explicitly constructed on the foundation of open scholarly infrastructure—using Crossref DOIs, ORCID researcher identifiers and open metadata as its backbone, augmented by AI-driven entity resolution, classification and data enrichment [27]. The central thesis of its design was that by linking publications, grants, clinical trials, patents and policy documents into a single interlinked graph, it could provide a broader context for research than the traditional publication-citation ecosystem alone [27, 26].

From its inception, Dimensions pursued a strategy of progressive data democratisation. A free version was made available in 2018 on the principle that researchers should be able to search the scholarly record without charge, and that analyses used in research evaluation should be reproducible against accessible data [24]. Subsequently, the full dataset was made available on Google BigQuery, democratising not only access to data but access to the computational capacity required to analyse it at scale [28]. The rationale was that the combination of accessible data and on-demand computation could lower barriers for researchers, analysts and policymakers who had previously been excluded from large-scale bibliometric analysis by the cost of both data and infrastructure.

Each of these steps moved the inner boundary of the annulus outward for the academic research community. But Dimensions also maintains commercial products built on higher-order data refinement—institutional analytics, research landscape mapping, funder intelligence, patent-publication linkage—that serve the differentiated demand described in Sec. III. The same underlying scholarly record serves radically different user communities at different quality levels. The commercial revenue from frontier refinement supports the infrastructure costs of maintaining the open and freely available layers.

An earlier episode in Digital Science’s history illustrates the annulus lifecycle for a single data type with particular clarity. In 2015, Digital Science created the Global Research Identifier Database (GRID)—a comprehensive, curated database of research organisations worldwide, designed to solve the institutional disambiguation problem that the paper identifies as a key production friction. GRID was built because no open, community-governed organisational identifier system existed at the time, and Dimensions needed reliable institutional resolution as part of its data backbone. In December 2016, Digital Science released GRID under a Creative Commons CC0 licence—placing the entire dataset in the public domain without restriction. GRID was subsequently made available as Linked Open Data [64] and grew to cover over 100,000 institutions. Then in 2021, once the Research Organization Registry (ROR) had built sufficient community support, coverage, and governance maturity, Digital Science discontinued public releases of GRID in favour of ROR—effectively passing the torch to a community-governed identifier system that GRID had helped to catalyse. This trajectory traces the complete annulus lifecycle for organisational identifiers within a single organisation’s experience: Frontier investment (creating GRID), deliberate intervention to push the inner boundary outward (CC0 release), and eventual transition to community governance (ROR) when the community was ready to sustain it. It also illustrates the structural constraint that open provision depends on prior frontier investment: ROR exists in its current form because someone first bore the cost of building the comprehensive organisational dataset from which the community effort could grow.

This practical experience illustrates several of this paper’s theoretical claims. First, that the annulus has sectoral dimensions: different data types and different quality levels occupy different positions, and the inner boundary moves outward at different rates for different user communities. In terms of Sec. III.4, Dimensions’ trajectory represents a case where the inner arc has been deliberately pushed further out than market forces alone would produce—a strategic choice to increase the openness ratio for academic metadata in the annulus. Second, that the annulus is shaped by technological capability: as AI-enabled structuring lowers the cost of basic normalisation, data types that were previously commercially valuable become candidates for open provision. Third, that the boundary between open and commercial is not static but is determined dynamically by the interaction of production cost and user demand—and that a commercial actor can, by strategic choice, accelerate the outward movement of the inner boundary. Fourth, that the constraint identified in Sec. III.4—that the inner boundary cannot move outward faster than the outer boundary—operates in practice: the open layers of Dimensions are sustained, in part, by the investment in frontier refinement that drives the outer boundary forward.

The limitations of this model should also be noted. The governance of the boundary between open and commercial layers is determined by the commercial actor rather than by community oversight. Open provision is contingent on the commercial operation remaining viable. And the reliance on a single commercial actor for a significant portion of open data infrastructure creates concentration risks. These limitations reinforce this paper’s argument that governance norms of the kind the Barcelona Declaration is beginning to develop are needed to ensure that the boundary is managed in the public interest and not according to commercial logic.

IX Open Questions

The framework developed in this paper raises several questions that define a research agenda at the intersection of science policy, innovation economics and knowledge governance.

Optimal annulus width. The welfare framework in Sec. IV yields qualitative predictions about how the optimal annulus width should vary by data type, but making these predictions quantitative requires empirical estimates of the functional forms: the social benefit function B​(ri), the production cost function C​(ro), the maintenance cost M​(ri), and the revenue function R​(w). None of these is currently known for any data type in the scholarly ecosystem, and their estimation presents significant methodological challenges—not least because the “radial” dimension of the annulus is not directly measured in any existing dataset. Developing proxy measures—for example, using the time lag between commercial and open availability of specific metadata types as a proxy for annulus width, or using infrastructure provider cost data to estimate M—would ground the framework quantitatively. The dynamic extension, casting the problem as an optimal control problem with technological change driving the evolution of both boundaries, would connect to the literature on optimal technology diffusion and would yield predictions about the trajectory of annulus compression that could be tested using longitudinal data.

Efficiency measurement and the openness ratio. Can the distance between the actual scholarly data production system and the efficient benchmark be measured? The openness ratio introduced in Sec. III.4—the ratio of inner to outer radius for each data-type segment—provides one candidate metric. Tracking this ratio over time for specific data types (citation data, institutional affiliation, CRediT metadata, funding acknowledgements) would yield an empirical measure of how quickly the open core is expanding relative to frontier structuring. Additional proxies include the proportion of metadata fields that require downstream normalisation, the rate at which standards adoption reduces processing costs, and the volume of duplicated AI-derived enhancement across the system. Decomposing annulus width into its technical and legal components, as discussed in Sec. II, would further refine the analysis.

Governance mechanism for the inner boundary. What institutional mechanisms could ensure the inner boundary moves outward appropriately? The Crossref mechanism works for data types already in the publishing workflow. The Barcelona Declaration provides a normative framework, but Ostrom’s [51] design principles predict that governance without effective monitoring and graduated sanctions will erode under pressure. Empirical study of signatory behaviour change—analogous to the compliance studies conducted for DORA [70]—would test the Declaration’s effectiveness.

Metadata provenance standards. The distinction between principled and performative openness requires a formal metadata provenance standard. Existing foundations—the W3C PROV standard, DataCite’s provenance model, Crossref’s governance framework—could serve as building blocks. A standard analogous to verified carbon offsets would provide a more transparent basis for assessing the reliability of structured metadata.

Equity implications. How does the current annulus structure affect institutions in under-resourced settings? What level of open baseline provision is required to ensure that all institutions can participate meaningfully in the research ecosystem? These are empirical questions that could be addressed through comparative studies of data access and analytical capability across institutions of different resource levels.

Cross-domain generalisability. Similar annulus dynamics appear in clinical data, environmental monitoring, geospatial data commons and AI training dataset governance. Comparative analysis would test the model’s generality and refine its parameters.

X Conclusion

The persistent framing of scholarly knowledge infrastructure as a contest between openness and commerce has obscured a more productive question: how should the boundary between open and commercially refined data be governed so that the system as a whole serves the public interest?

This paper has argued that the innovation annulus—the zone between the open core and the advancing frontier of refined knowledge products—is not a pathology but a functional and permanent feature of the knowledge ecosystem. It exists because the cost of producing and refining structured knowledge data is real and persistent, shaped by production frictions that technology reduces but cannot eliminate. It is sustained by differentiated demand from communities—competitive institutions, research translation partnerships, corporate R&D—whose quality requirements exceed what open provision can deliver. And its width for any given data type is a measure of the system’s distance from perfect production efficiency: a distance that shrinks as standards mature and technology advances, but that is continuously regenerated as new data types emerge at the frontier.

Although we have principally confined our attention to consideration of scholarly metadata, we believe that this model can also be useful in thinking about the evolution of open access, and the commercial software / open source ecosystem where, with the rise of vibe coding, the following comments on AI are particularly pertinent.

AI has changed the parameters of this system profoundly. It has lowered the cost of basic structuring, pushing the inner boundary outward and expanding the volume of data that can be provided openly. But it has also raised the quality threshold at which refinement has value, shifted the commercially relevant frontier to higher-order tasks, while simultaneously introducing systemic risks through unprovenanced AI-derived metadata. Again, the annulus persists—not through resistance to openness, but because the economics of data refinement in a system with technological frictions and differentiated demand make it a structural feature of the landscape.

The governance question is how to ensure the annulus is the right width. The efficient market analogy suggests a benchmark: the annulus should reflect genuine production costs and nothing more. The welfare framework developed in Sec. IV makes this more precise: the optimal width for any given data type is determined by the interaction of the social benefit of openness, the cost of frontier production, the cost of maintaining the open core, and the financial sustainability of the system as a whole. Where the annulus has historically been inflated by market power and institutional lock-in—as in the Web of Science / Scopus duopoly era—competition and standards adoption have compressed it toward a more efficient width. Where it persists at the frontier—in domain-specific enrichment, integrity signals, and emerging metadata types—it reflects genuine investment that the system needs someone to make. Critically, the framework shows that an annulus compressed below the sustainable level does not merely reduce commercial revenue—it slows the advance of the outer boundary and, with it, the growth of the open core.

The Barcelona Declaration, with its framework of research information citizenship and its constituency of funders, institutions and infrastructure providers, represents the most promising forum for developing the community norms that should govern the inner boundary. It could, with appropriately expanded governance, become the venue where the scholarly community develops a shared understanding of which data types should be inside the open core, what quality and provenance standards apply, and what responsibilities producers, consumers and aggregators of metadata owe to each other. Whether it will develop the institutional machinery to make this vision durable—monitoring, enforcement, conflict resolution—remains to be seen.

The annulus will persist—the width is the price of being able to trust the data to the level needed to meet a given use case. The question is whether we govern it wisely: ensuring that the inner boundary moves outward as technology matures, that the open core is built on principled rather than performative foundations, and that the scholarly data ecosystem serves all its users—including those who cannot afford to pay for frontier refinement—as equitably and efficiently as the state of the art allows.

Acknowledgements.

The author wishes to thank Mark Hahnel for valuable suggestions around Sec. IV, and Bianca Kramer, Simon Porter, Ludo Waltman, and Juergen Wastl their careful reading of this manuscript. Any remaining errors or omissions lay solely with the author. Conflict of Interest Statement. The author is CEO of Digital Science, a technology company that operates within the scholarly knowledge infrastructure landscape analysed in this paper. Digital Science’s portfolio includes Dimensions, a bibliographic database that occupies a position within the annulus as defined herein, as well as Altmetric and other research analytics products. The author’s position within this landscape informed the analysis but also constitutes a potential conflict of interest: the framework developed here could be read as a justification for the business model of the author’s employer. We have sought to present the structural dynamics as objectively as possible, and we note that practitioner knowledge of the actual economics of data production is precisely what has been missing from much of the largely academic literature on open infrastructure. Readers should nonetheless be aware of this positionality when evaluating the arguments presented.

References

[1] J. Adams, K. Gurney, D. W. Hook, and L. Leydesdorff (2014) International collaboration clusters in Africa. Scientometrics 98 (1), pp. 547–556. https://dx.doi.org/10.1007/s11192-013-1060-2

[2] J. Adams and M. Szomszor (2024) National research impact is driven by global collaboration, not rising performance. Scientometrics 129 (5), pp. 2883–2896. https://dx.doi.org/10.1007/s11192-024-05010-6

[3] J. Adams (2013) The fourth age of research. Nature 497 (7451), pp. 557–560. https://dx.doi.org/10.1038/497557a

[4] L. Allen, V. Kiermer, S. Porter, and R. Whittam (2025) A ten-year drive to credit authors for their work—and why there’s still more to do. Nature 648 (8092), pp. 33–34. https://dx.doi.org/10.1038/d41586-025-03860-5

[5] L. Allen, A. O’Connell, and V. Kiermer (2019) How can we ensure visibility and diversity in research contributions? How the Contributor Role Taxonomy (CRediT) is helping the shift from authorship to contributorship. Learned Publishing 32. https://dx.doi.org/10.1002/leap.1210

[6] L. Allen, J. Scott, A. Brand, M. Hlava, and M. Altman (2014) Publishing: credit where credit is due. Nature 508 (7496), pp. 312–313. https://dx.doi.org/10.1038/508312a

[7] K. Arrow (1962) Economic welfare and the allocation of resources for invention. The Rate and Direction of Inventive Activity: Economic and Social Factors, pp. 609–626.

[8] A. B. Atkinson and J. E. Stiglitz (1980) Lectures on public economics. McGraw-Hill, New York.

[9] Barcelona Declaration on Open Research Information (2024) Barcelona declaration on open research information. https://barcelona-declaration.org

[10] C. L. Borgman (2015) Big data, little data, no data: scholarship in the networked world. MIT Press.

[11] S. C. Bradford (1934) Sources of information on specific subjects. Engineering 137, pp. 85–86.

[12] A. Brand, L. Allen, M. Altman, M. Hlava, and J. Scott (2015) Beyond authorship: attribution, contribution, collaboration, and credit. Learned Publishing 28 (2), pp. 151–155. https://dx.doi.org/10.1087/20150211

[13] L. Butler, L. Matthias, M. Simard, P. Mongeon, and S. Haustein (2023) The oligopoly’s shift to open access: how the big five academic publishers profit from article processing charges. Quantitative Science Studies 4 (4), pp. 778–799. https://dx.doi.org/10.1162/qss%5Fa%5F00272

[14] G. Colavizza, I. Hrynaszkiewicz, I. Staden, K. Whitaker, and B. McGillivray (2020) The citation advantage of linking publications to research data. PLoS ONE 15 (4), pp. e0230416. https://dx.doi.org/10.1371/journal.pone.0230416

[15] G. Czépán and A. Dima (2024) Representation of the Global South in bibliometric databases: a systematic review. Scientometrics 129, pp. 2109–2131.

[16] P. Dasgupta and P. A. David (1994) Toward a new economics of science. Research Policy 23 (5), pp. 487–521. https://dx.doi.org/10.1016/0048-7333%2894%2901002-1

[17] H. Debat and D. Babini (2020) Plan S in Latin America: a precautionary note. Scholarly and Research Communication 11 (1). https://dx.doi.org/10.22230/src.2020v11n1a347

[18] P. N. Edwards, S. J. Jackson, M. K. Chalmers, G. C. Bowker, C. L. Borgman, D. Ribes, M. Burton, and S. Calvert (2013) Knowledge infrastructures: intellectual frameworks and research challenges. Technical report Deep Blue, University of Michigan. https://dx.doi.org/10.3998/3336451.0014.101

[19] E. F. Fama (1970) Efficient capital markets: a review of theory and empirical work. The Journal of Finance 25 (2), pp. 383–417.

[20] N. T. Gallini (1992) Patent policy and costly imitation. RAND Journal of Economics 23 (1), pp. 52–63.

[21] E. Garfield (1980) Bradford’s law and related statistical patterns. Current Contents (19), pp. 5–12. Note: Reprinted in Essays of an Information Scientist, vol. 4, pp. 476–483

[22] T. Gillett (2021) Relaunching the academic seminar. https://www.researchinformation.info/interview/relaunching-academic-seminar-0/

[23] G. Hendricks (2023) Crossref and responsible AI data use. https://www.crossref.org/blog/

[24] C. Herzog, D. Hook, and S. Konkiel (2020) Dimensions: bringing down barriers between scientometricians and data. Quantitative Science Studies 1 (1), pp. 387–395. https://dx.doi.org/10.1162/qss%5Fa%5F00020

[25] D. Hicks, P. Wouters, L. Waltman, S. de Rijcke, and I. Rafols (2015) Bibliometrics: the Leiden Manifesto for research metrics. Nature 520 (7548), pp. 429–431. https://dx.doi.org/10.1038/520429a

[26] D. W. Hook, S. J. Porter, H. Draux, and C. T. Herzog (2021) Real-time bibliometrics: Dimensions as a resource for analysing aspects of COVID-19. Frontiers in Research Metrics and Analytics 5, pp. 595299. https://dx.doi.org/10.3389/frma.2020.595299

[27] D. W. Hook, S. J. Porter, and C. Herzog (2018) Dimensions: building context for search and evaluation. Frontiers in Research Metrics and Analytics 3, pp. 23. https://dx.doi.org/10.3389/frma.2018.00023

[28] D. W. Hook and S. J. Porter (2021) Scaling scientometrics: Dimensions on Google BigQuery as an infrastructure for large-scale analysis. Frontiers in Research Metrics and Analytics 6, pp. 656233. https://dx.doi.org/10.3389/frma.2021.656233

[29] D. W. Hook (2024) Barcelona: a beautiful horizon. https://www.digital-science.com/blog/2024/05/barcelona-a-beautiful-horizon/

[30] I. Hrynaszkiewicz, A. Birukou, M. Astell, S. Swaminathan, A. Kenall, and V. Khodiyar (2017) Standardising and harmonising research data policy in scholarly publishing. International Journal of Digital Curation 12 (1), pp. 65–71. https://dx.doi.org/10.2218/ijdc.v12i1.531

[31] A. B. Jaffe and J. Lerner (2004) Innovation and its discontents: how our broken patent system is endangering innovation and progress, and what to do about it. Princeton University Press.

[32] A. B. Jaffe (1986) Technological opportunity and spillovers of R&D: evidence from firms’ patents, profits, and market value. American Economic Review 76 (5), pp. 984–1001.

[33] S. Jasanoff (2005) Designs on nature: science and democracy in Europe and the United States. Princeton University Press.

[34] B. Z. Khan (2008) An economic history of copyright in Europe and the United States. In EH.Net Encyclopedia, R. Whaples (Ed.), https://eh.net/encyclopedia/an-economic-history-of-copyright-in-europe-and-the-united-states/

[35] S. Y. Khoo (2019) Article processing charge hyperinflation and price insensitivity: an open access sequel to the serials crisis. LIBER Quarterly 29 (1), pp. 1–18. https://dx.doi.org/10.18352/lq.10280

[36] V. Kiermer, A. Mudditt, and N. O’Connor (2025) Rethinking how we publish to support open science. Learned Publishing. https://dx.doi.org/10.1002/leap.2006

[37] P. Klemperer (1990) How broad should the scope of patent protection be?. RAND Journal of Economics 21 (1), pp. 113–130.

[38] B. Kramer, C. Neylon, and L. Waltman (2024) Barcelona declaration on open research information. https://dx.doi.org/10.5281/zenodo.10958522

[39] J. I. Lane and N. Potok (2024) Democratizing data: our vision. Harvard Data Science Review Special Issue 4. https://dx.doi.org/10.1162/99608f92.03719804

[40] J. I. Lane (2020) Democratizing our data: a manifesto. MIT Press.

[41] J. Lane, J. Owen-Smith, and B. A. Weinberg (2024) How to track the economic impact of public investments in AI. Nature 630 (8016), pp. 302–304. https://dx.doi.org/10.1038/d41586-024-01721-1

[42] V. Larivière, S. Haustein, and P. Mongeon (2015) The oligopoly of academic publishers in the digital era. PLoS ONE 10 (6), pp. e0127502. https://dx.doi.org/10.1371/journal.pone.0127502

[43] M. Mazzucato (2013) The entrepreneurial state: debunking public vs. private sector myths. Anthem Press.

[44] M. Mazzucato (2018) Mission-oriented innovation policies: challenges and opportunities. Industrial and Corporate Change 27 (5), pp. 803–815. https://dx.doi.org/10.1093/icc/dty034

[45] M. Mazzucato (2018) The value of everything: making and taking in the global economy. Allen Lane.

[46] M. K. McNutt, M. Bradford, J. M. Drazen, B. Hanson, B. Howard, K. H. Jamieson, V. Kiermer, E. Marcus, B. K. Pope, R. Schekman, S. Swaminathan, P. J. Stang, and I. M. Verma (2018) Transparency in authors’ contributions and responsibilities to promote integrity in scientific publication. Proceedings of the National Academy of Sciences 115 (11), pp. 2557–2560. https://dx.doi.org/10.1073/pnas.1715374115

[47] D. Mills (2024) One index, two publishers and the global research economy. Oxford Review of Education, pp. 1–16. https://dx.doi.org/10.1080/03054985.2024.2348448

[48] A. Mingardi (2015) A critique of Mazzucato’s Entrepreneurial State. Cato Journal 35 (3), pp. 603–625.

[49] P. Mongeon and A. Paul-Hus (2016) The journal coverage of Web of Science and Scopus: a comparative analysis. Scientometrics 106 (1), pp. 213–228. https://dx.doi.org/10.1007/s11192-015-1765-5

[50] W. D. Nordhaus (1969) Invention, growth, and welfare: a theoretical treatment of technological change. MIT Press.

[51] E. Ostrom (1990) Governing the commons: the evolution of institutions for collective action. Cambridge University Press.

[52] C. D. Peeler (1999) From the providence of kings to copyrighted things (and French moral rights). Indiana International and Comparative Law Review 9 (2), pp. 423–456.

[53] S. Pinfield, J. Salter, and P. A. Bath (2016) The “total cost of publication” in a hybrid open-access environment: Institutional approaches to funding journal article-processing charges in combination with subscriptions. Journal of the Association for Information Science and Technology (67), pp. 1751–1766. https://dx.doi.org/10.1002/asi.23446

[54] S. Pinfield (2024) Achieving global open access: The need for scientific, epistemic and participatory openness. Routledge, London. External Links: ISBN 9781032679259

[55] S. J. Porter and D. W. Hook (2022) Connecting scientometrics: Dimensions as a route to broadening context for analyses. Frontiers in Research Metrics and Analytics 7, pp. 835139. https://dx.doi.org/10.3389/frma.2022.835139

[56] S. Porter (2024) The Barcelona Declaration: exploring our responsibilities as metadata consumers. https://www.digital-science.com/blog/2024/07/

[57] S. Porter (2026) No shortcuts to research information citizenship. https://www.digital-science.com/blog/2026/02/

[58] J. Priem, H. Piwowar, and R. Orr (2022) OpenAlex: a fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv preprint arXiv:2205.01833.

[59] S. Scotchmer (2004) Innovation and incentives. MIT Press.

[60] C. Shapiro and H. R. Varian (1999) Information rules: a strategic guide to the network economy. Harvard Business School Press.

[61] J. E. Stiglitz (2000) Economics of the public sector. 3rd edition, W. W. Norton.

[62] D. E. Stokes (1997) Pasteur’s quadrant: basic science and technological innovation. Brookings Institution Press.

[63] P. Suber (2012) Open access. MIT Press.

[64] M. Szomszor, J. Adams, R. Fry, C. Gebert, D. A. Pendlebury, R. W. K. Potter, and G. Rogers (2020) Interpreting bibliometric data. Frontiers in Research Metrics and Analytics 5, pp. 628703. https://dx.doi.org/10.3389/frma.2020.628703

[65] J.A. Teixeira da Silva and S. Nazarovets (2022) The role of publons in the context of open peer review. Publishing Research Quarterly (38), pp. 760–781. https://dx.doi.org/10.1007/s12109-022-09914-0

[66] P. Vincent (2021) Rethinking the research seminar for a post-COVID world with cassyni. https://blogs.lse.ac.uk/impactofsocialsciences/2021/09/01/rethinking-the-research-seminar-for-a-post-covid-world-with-cassyni/

[67] L. Waltman (2020) Responsible research assessment requires open scholarly metadata. https://dx.doi.org/10.5281/zenodo.4021492

[68] N. Wiener (1954) The human use of human beings: cybernetics and society. 2 edition, Doubleday Anchor, Garden City, NY.

[69] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, et al. (2016) The FAIR guiding principles for scientific data management and stewardship. Scientific Data 3, pp. 160018. https://dx.doi.org/10.1038/sdata.2016.18

[70] P. Wouters, C. R. Sugimoto, V. Larivière, M. E. McVeigh, B. Pulverer, S. de Rijcke, and L. Waltman (2019) Rethinking impact factors: better ways to judge a journal. Nature 569, pp. 621–623. https://dx.doi.org/10.1038/d41586-019-01643-3

[71] P. Wouters (1999) The citation culture. Ph.D. Thesis, University of Amsterdam, Amsterdam. https://hdl.handle.net/11245/1.163066

Editors

Kathryn Zeiler
Editor-in-Chief

Jason Chin
Handling Editor

Editorial assessment

by Jason Chin

DOI: 10.70744/MetaROR.427.1.ea

The three reviewers regard the paper as a valuable and timely contribution, crediting it with supplying vocabulary and analytical tools for a debate that has largely proceeded without either. They all also identify the governance material as its strongest element. Their shared criticism is that the paper’s main claim, that the innovation annulus is a permanent structural feature, follows from the construction of the model rather than from evidence about the system it describes. The reviewers also ask for closer engagement with institutional and financial models that structure and enrich metadata without enclosing it (OpenAlex, OpenAIRE, EuropePMC, commons-based governance and diamond arrangements). Further, they question the treatment of production friction as exogenous when incumbents’ strategic choices plainly shape where that friction sits and who bears it. The discussion around the framework, and any future revisions, should address distributional consequences for actors who cannot pay.

Recommendations for enhanced transparency

  • Add author ORCID iD.
  • Add a Data Availability Statement to report that no data are used in the article.
  • Add a funding source statement. Authors should report all funding in support of the research presented in the article. Grant reference numbers should be included. If no funding sources exist, explicitly state this in the article.

For more information on these recommendations, please refer to our author guidelines.

Competing interests: None.

Peer review 1

Bianca Kramer

DOI: 10.70744/MetaROR.427.1.rv1

“Other models are available”

Note on positionality: the author of this response is independent advisort and research analyst at Sesame Open Science (working in the areas of open science, open metadata and open infrastructure) as well as Executive Director of the Barcelona Declaration on Open Research Information. This response is written in a personal capacity, representing my personal views and opinions.

The paper introduces and discusses the ‘innovation annulus’ as a zone of closed structured metadata that separates a core of fully open metadata and an advancing frontier of refined knowledge products, and argues that the annulus exists because the cost of producing and refining structured knowledge data is real and persistent, shaped by production frictions that technology reduces but cannot eliminate.

In this review, I focus on a number of counterarguments to the premise of the article. The review does not go into detail on the mathematical details of the model presented, but instead, hopes to contribute to a discussion on the assumptions underlying the model as a whole.

From zero-sum game to win-win scenario?

The author argues that the debate about scholarly knowledge infrastructure has traditionally been framed as a zero-sum game between openness and commercial enclosure, with every advance in openness a retreat for commercial interests and vice versa.

The proposed model is presented as a positive sum game (a win-win scenario) where a base layer of structured metadata is openly available, and development of new or enriched metadata is in the hands of commercial providers. The paper further sees a role for both economics and community governance in determining where the dividing line between the two classes of metadata should lie (recognizing that this can differ for different types of metadata).

This model keeps commercial interests at its center by postulating that innovation can or will only take place in a commercial setting. In effect, this perpetuates a dependency on commercial systems, not only as a locus of innovation, but also as source of structured and refined metadata that kept close (for now) to recoup financial investments and make a profit.

Other models are available

In the paper, the existence of a zone of closed structured metadata is justified by stating that the cost of producing and refining structured knowledge data is real and persistent. In our view, the latter is a given, but the conclusions derived from that in the paper are not.

Provision of structured metadata at source

First, producers of scholarly metadata play an important role in providing structured metadata at source. An obvious example is publishers depositing publication metadata through Crossref, but this also involves institutional and subject repositories that expose metadata for publications, as well as data repositories and software repositories.

The author acknowledges that provision of better structured metadata at source (helped by technical advances, standardization and community norms) does reduce third-party efforts for metadata structuring and enrichment. This increases the proportion of metadata (of a given type) that is openly available, reducing the width of zone of closed structured metadata. There are, however, potential additional dynamics at play that could influence this provision of open metadata at source. When publishers provide access to full text or JATS XML access to bibliographic databases to use in the extraction and structuring of metadata, there may be less incentive to provide those same metadata openly at source. A similar development has been observed with publishers requesting (open) bibliographic databases to take down abstracts at a time where abstracts are increasingly valuable as training. material for LLMs (Kramer, B., 2024 & Tay, A., 2025).. In these cases, the legal and contractual frictions described by the author may contribute to a non-level playing field for metadata structuring and enrichment, and the existence of a closed zone of structured metadata in itself may limit the provision of structured metadata at source.

Alternative financial models for structuring and enriching metadata

Second, where third-party efforts are required to harmonize, structure and enhance scholarly metadata, there are multiple examples of this work being taken on not as commercial activity (with the resulting metadata, at least initially, being kept in the closed zone), but by organizations that operate under different financial models and, make the resulting metadata immediately available as open metadata as part of their ethos and practice. While the author discussed a limited role for ‘state investment’, the financial models used by infrastructures that provide the metadata they enrich and structure directly as open metadata are much more varied than this term suggests.

For example, OpenAlex receives project funding from charitable funders for innovation, but also financial support from research performing organizations and funders through their institutional membership route, and direct revenue for services provided on top of their database of openly available metadata. OpenAIRE, originally a direct recipient of European Commission funding through consecutive Framework Programmes, has diversified revenue streams through an institutional membership programme, participation in funded projects and direct collaborations with e.g. national governments and library consortia. As a third example, EuropePMC works on longer-term operational funding from a group of both national and charitable funders in the medical domain to link and enrich metadata in a specialized domain, going far beyond publication metadata only.

Crucially, all these models involve decisions by research institutions and funders to financially contribute to the generation of structured enriched metadata that is then made openly available rather than kept closed to be licensed to other users.

A different look at the frontier

The paper argues that more specialized demands for metadata, e.g. for contextualised analytical products, data on research capabilities beyond publication counts and (often domain-specific) data to support corporate R&D activities, require development and innovation by commercial actors, as they “require a quality level that raw open metadata does not yet consistently deliver”. A counterargument to this would be that, like above, there is no inherent reason why such development and innovation cannot be supported by other financial models, if downstream users would elect to pay for these. The difference is not raw open data versus closed high-quality data, but high-quality data directly released as open data or restricted as closed data. An underlying question could be whether corporate research organizations would consider it in their (competitive) interest to not just pay for access to data, but for the creation of high-quality data that would also be publicly available. Here, it should be noted that any competitive advantage from paying for closed data would be mitigated by the fact that the same data would be available to any other party willing and able to pay for them. There is an additional argument against considering work on ‘new’ metadata types as naturally in the purview of commercial development because these types of data only serve specialized usage. This is that limited, closed availability of these data types in itself slows down uptake and usage, and by definition excludes lesser resources actors, not because they have no use for these data types, but because they cannot afford access to them. Of the examples given in the paper, funding flow analysis stands out as an area where there is a lot of interest from research funders in the Global South (see https://www.clacso.org/fundingflows/ and https://idrc-crdi.ca/en/what-we-do/projects-we-support/project/state-science-technology-and-innovation-africa-science. Here it should also be noted that access to otherwise closed data for specific users (e.g. in the context of a research project), especially without the right to share the data, does not represent the same value and benefits as true open availability of such data does.

Finally, when considering the ‘frontier’ of specialized metadata and metadata usage, a distinction can be made between the data itself and (analytical or other) applications and services built on top of these data. It could be argued that a financial model that charges for services while having the underlying data openly available would be in line with the Principles of Open Scholarly Infrastructure, at least for this aspect. In addition, it would keep the field open for (competitive) innovation and development to take place on top of open metadata.

The role of community

The paper sets out a role for the scholarly community, and explicitly the Barcelona Declaration, to develop a normative shared understanding of which data types should be inside the ‘open core’, what quality and provenance standards should apply, and what responsibilities producers, consumers and aggregators of metadata should have towards each other. While there certainly is value in such collaborative discussions, they should not be positioned to implicitly endorse a model where the existence of a zone of closed structured metadata, produced and restricted by commercial actors, is considered both inevitable and inherently beneficial.

As the closing sentence of the paper reads: “The question is whether we govern it wisely: ensuring that (…) the scholarly data ecosystem serves all its users—including those who cannot afford to pay for frontier refinement— as equitably and efficiently as the state of the art allows.” This ambition deserves the consideration of multiple models for enabling the structuring and enrichment of metadata that truly benefit all users, as well as the role that individual research performing and funding organizations have in deciding where to allocate both financial resources and in kind efforts (e.g. in participating in governance bodies and the integration of data sources in institutional processes).

This is not to discount the potential value of commercial actors in this space (especially when they operate on a service- not data-based revenue model), but to challenge their ‘natural’ role in the provision of high-quality metadata. The argument is not about whether producing metadata is somehow easy or cost-free (it isn’t), but about the different ways this production can be organized and financed.

Final remark

The paper offers a valuable contribution in theorizing and modelling the forces at play in shaping the way scholarly metadata are created, structured, enriched and made available for use by the scholarly community. Further explorations of the model under different assumptions, including those outlined in this response, could contribute to the discussion of the role of different types of actors (including commercial actors) in this space.

Factual corrections

Crossref

The authors state “Crossref has an increasingly diverse membership including publishers, research institutions, funders and governmental organisations.” – this would benefit from the clarification that all Crossref members are organizations that register DOIs for content items – which can indeed include organizations with institutional publishing activities and funders registering grants.

Barcelona Declaration

The author states “The Declaration’s membership includes funders, institutions, and infrastructure providers” – to clarify, the Barcelona Declaration does not not have
membership, but rather signatories and supporters, which are two distinct categories. The commitments of the Barcelona Declaration are aimed at organizations performing, funding and evaluating research, and these types of organizations can become signatories of the Declaration. Organizations providing services, data and infrastructure around open research information can be considered as supporters. Overall, the Barcelona Declaration envisions changing institutional practices through supporting internal processes at institutions, broadened advocacy and collective action.

Competing interests: The author of this response is independent advisor and research analyst at Sesame Open Science (working in the areas of open science, open metadata and open infrastructure) as well as Executive Director of the Barcelona Declaration on Open Research Information. This response is written in a personal capacity, representing my personal views and opinions.

Peer review 2

Cameron Neylon

DOI: 10.70744/MetaROR.427.1.rv2

This manuscript presents a theoretical framework drawn from efficient markets analysis and a parallel argument for the role of commercial innovation in the production of open scholarly metadata. It presents several valuable insights and tools for examining the economics of scholarly metadata provision and makes a strong argument for the importance of purposefully designed governance in driving the development of openness (and identifying where it is uneconomic).

In this review I take as a rhetorical goal of the paper making an argument for identifying the role of commercial innovation and capital in the optimal production of scholarly metadata or research information. In particular I am working from the perspective that the goal is to make a case to those sceptical of the value and role of commercial players and external capital.

As a consequence I should declare my own priors. I am generally sympathetic to the view that commercial actors should not be automatically excluded in principle from the community of open research information production. However, I am sceptical in practice that commercial actors can be constructed with appropriate incentives and governance safeguards. In that sense I believe I might be considered a reasonable example of the target audience.

The paper provides a wealth of valuable insights and analytical frames. However, I believe it fails in its rhetorical goals partly due to faults in its formal analysis, and partly due to the limitations of formal arguments in persuasion. In this review I want to start with the rhetorical issues prior to specific criticisms of the argument.

Rhetorical structure of the paper

The paper proceeds from a formal argument based on a model, proceeds to develop an analytical framework which expands on the formal model, and then applies this in general terms to use cases and argumentation about how to move forward. The challenge with the structure is that by starting with a very strong claim built on a formal model the rhetorical structure, particularly for a sceptical target audience, is weak. In common with all such economic models they reproduce their own assumptions and fail to capture important complexities of the underlying systems. In turn these complexities are what drive the actual outcomes. A classical example of this, very relevant to the current paper is Ostrom’s dissection of Hardin’s Tragedy of the Commons.

Reading from the front, the sceptical reader will therefore seek to identify issues with the formal argument. Inevitably, due to the nature of formal arguments, these will be found and the sceptical reader is therefore unpersuaded. In contrast, the sympathetic reader will agree with the outline of the formal argument and proceed to the analysis. Anecdotally this aligns with the reception of the preprint that I have observed.

However, if the paper is “read in reverse” it becomes substantially more persuasive. Starting from the strong point on governance, it proceeds to develop some examples of that governance in practice. This includes some novel – even startling – insights into existing systems, and might merit further development, here or elsewhere. For instance, the notion of Crossref’s evolution as an “openness ratchet” has some potential alignments with considerations of the evolution of club-like economic structures and addresses aspects of the welfare-investment tradeoff in a way that seem deserving of further attention.

With the governance and examples in hand, the analytic value of the model is clear – not as an argument of the inevitability of the annulus but as an analysis of its characteristics in practice. In my view the argument that in practice there is fairly strong co-alignment of limitations on openness and modes of ensuring return on capital investment makes the analytical approach valuable. However, this does not strengthen the case for this being a necessary condition of the system, but merely a necessary consequence of the construction of the model. My view is that the value of the paper, particularly for the sceptical reader, is reduced by the strong leading claims, rather than working towards the conditions of value creation in pragmatic terms.

More crudely, as currently structured, the paper reads to the unsympathetic reader as a strong claim for the necessity of including commercial innovation and capital in community systems, rather than an analysis of where boundaries might be placed for maximum welfare. Read from end to beginning the analytical value is clearer and the argument for applying this analysis in considerations of strategy and governance is clearer.

Formal argument

As noted, the paper leads with a strong claim that the annulus is a necessary condition of the structuring of metadata. The claim rests on an implicit model that boundaries of innovation, structure and openness are all co-incident (or near co-incident). For instance in the legend to Figure 1 “The width of the annulus at any point represents the gap between the current frontier of commercially refined data and the current baseline of open provision” [emphasis added] conflates two issues in a way crucial to the structuring of the model but which need not be connected in theory, even if the argument is that they often are in practice.

From a purely mathematical perspective these classes of arguments can collapse under conditions of high dimensionality, uneven or heterogenous boundaries and other topologies (e.g. an inverse model where pockets of “unstructured” data exist separately within a universe of “open metadata”). These purely formal issues can relate to issues of interpretation of the model so they can be of value to consider. One example of this might be boundary heterogeneity relating to differential subsidies across the boundary (e.g. as noted in the paper, the basing of Dimensions on open research information products amounts to a community subsidy of innovation, subsidies can also operate in the opposite direction). A second example is how the assumption that all possible metadata is constructed in a connected field removes different economic models where there is not fungibility or arbitrage possible between them from consideration.

Another might be competition in the innovation space that differentially targets communities with differing norms. For example a corporate entity might intentionally generate open products with the goal of capturing scholarly markets, whereas a competitor may be either restrained from doing so, or focus on different markets where the same forms of openness are not valued as a market differentiator.

A further assumption is that information starts unstructured, explicitly noted as “Frontier data types” under the three types of production. This makes a good example of how the rhetorical issue can play out. The paper has already noted that much data starts as (implicitly) structured and becomes unstructured, after which it is necessary to “restructure” it. The costs of standardisation are community costs and conventional innovation and friction models fail to capture the systems of governance and economics required to address these. Read as a statement that standardisation requires investment, this point is robust, but in the context of making a claim for a necessary role of commercial innovation it becomes easy to pick holes.

More formally this point also relates to the co-incidence assumption that underpins the whole argument. Questions of what metadata are required or desirable sit in complex relations to the community consensus on centralised systems that deliver those data that are part of the consensus. An example of this is the differential assumptions relating to journal metadata and (scholarly) book metadata. The strong statement that this makes a boundary a “permanent structural feature” and not merely a “legacy of historical infrastructure debt” understates the complexity and overstates the degree to which the model captures the system. Inverted, the argument is stronger in my view. Given that there will in practice be changes and arguments over whether they should be standardised, the analytical framework gives a set of tools for considering value creation and welfare maximisation for competing claims about what should be a focus for standardisation and investment.

This can be framed more politically. The strong claim is that community failures will be addressed by market and capitalist logics and this creates the necessity for governance forums to reappropriate innovation into the community space (with the consequent outward payments for appropriate returns to capital). Framed inversely, where there is community dissensus on what should be standardised, commercial innovation will seek to fill this gap, applying capitalist and non-community logics, undermining potential collective value creation. The conclusion, that governance systems are required to address these failure modes, is the same in both cases although the focus of that governance may be different.

As a concrete example, consider part of the historical infrastructure debt at hand, the continuing failure of large-scale commercial products in the journal-submission ecosystem to address expressed market needs for retaining and validating the structure of submitted metadata. The collective economic incentives for providing validation and structuring at point of submission are very high, but for a variety of reasons, including near monopoly, rent seeking, and the complexity of supply chains these incentives are not transferred to the point of service provision. It would be interesting in the current section II to see some examples of these issues worked through in addition to those currently described.

Analytical model and other minor issues

The remaining issues related here are largely minor quibbles or areas that might be deserving of further analysis.

Section IV.1

Assumptions around the structure of the functions V and C are (I think?) potentially necessary for the mathematics to hold. This is outside my expertise but my understanding is that these functions need to be differentiable and concavity/convexity is important. Some exploration of how this plays out in practice and whether the assumptions of V being concave and C convex may be valuable. In particular both make assumptions about the homogeneity of the field, even while the analysis is explicitly applied to new classes of metadata at the frontier.

It is plausible to postulate substantial discontinuities in V for instance, where critical mass and interconnection of metadata types substantially changes the value proposition. What are the consequences of such (potentially non-differentiable) discontinuities? Similarly the assumption that C is increasing and convex may come under pressure if there is a separation of infrastructure and marginal costs. Moreover the assumption that “the easy structures tasks…are accomplished first…” leads to increasing costs would be a characteristic of a functioning market, but may not be the case for any specific form of structuring.

More broadly, assumptions of homogeneity are a limitation on the general analytical power of the model and this should be addressed specifically through examples.

Section IV.2

This section sits at the centre of my point about the rhetorical structure of the paper. The predictions here are both the most interesting part of the overall paper but also consequences of the structure of the model. Framing them more within the assumptions of the model might seem to make the argument weaker but I would argue it makes the usefulness of the model as a means of clarifying points of disagreement stronger.

As an example Prediction 5 is in some ways the central claim of the paper. It advances the argument that the author has made in other settings that commercial investment, with (or despite, or even because of) its attendant limitations, adds substantial value. In the current model, this is a direct consequence of the structure of the model. It may for instance break down in cases where there is cross talk between different sets of structuring processes that exist in a complex relationship of costs and underpinning value to each other.

Or to put it another way, it is dependent on homogeneity and simple topology of the model. If the boundary is non-homogeneous then the open core in some areas can advance ahead of the structuring frontier in others. This is actually common, with community and publicly subsidised efforts creating both technological advances and metadata resources that collapse the costs structuring in other spaces. Commercial innovation can also play this role of course.

Section VIII

A minor point. The text as written indirectly implies that the full Dimensions dataset is freely available on Google BigQuery through the concatenation of two sentences in paragraph 3 (“…A free version was made available…Subsequently the full dataset was made available on Google BigQuery…”). Given the nuances of availability and “openness” are central to the paper, being clear that these are two separate initiatives seems important.

The discussion of GRID might also make reference to the history of ORCID and the stepping back of Thomson-Reuters from a product oriented position to supporting a community initiative as a parallel. There seems much value in emphasising cases where “…the inner arc has been deliberately pushed further out…” more generally. This again is central to the argument being made here and more generally. The role for responsibly acting commercial players and capital is worth exploring and these examples help to make that case as well as to examine how these kinds of opportunities can be encouraged.

Competing interests: I declare no competing interests beyond the philosophical and political perspectives noted in the review.

Peer review 3

Neil Jacobs

DOI: 10.70744/MetaROR.427.1.rv3

Summary

This manuscript develops a conceptual framework, centred on the notion of an “innovation annulus”, to describe the relationship between open scholarly metadata and commercially refined data products. Drawing analogies from financial economics and intellectual property theory, the author argues that a zone between open and commercial provision is a permanent and functional feature of the scholarly knowledge ecosystem. The paper further proposes a welfare-theoretic framework and derives qualitative predictions about the “optimal” width of this annulus, with implications for the governance of research information.

The manuscript is ambitious and clearly written. It addresses an important and timely topic, and offers a unifying conceptual lens intended to bridge debates around open infrastructure, commercial data provision, and the impact of AI on metadata production. The integration of economic analogies, policy discussion, and empirical examples is intellectually engaging and the proposed conceptual framework has the potential to be of significant value.

However, the paper also raises a number of substantial concerns regarding its conceptual foundations, empirical grounding, and normative implications. These issues limit its current contribution.

Major Comments

1. Ambiguity in the epistemic status of the framework

The manuscript is presented as a theoretical contribution, but its status remains unclear. It appears to operate simultaneously as a descriptive model of the scholarly data ecosystem, an explanatory account of its dynamics, a normative framework for governance and a source of empirically testable predictions. Yet it is not clearly established how it should be evaluated. In particular:

  • The framework is not developed as a formally testable theory with clearly specified empirical indicators.
  • The derived “predictions” are largely qualitative and, in some cases, appear either self-evident or difficult to falsify.
  • Despite this, the paper proceeds to make substantive claims about policy and governance.

This raises a central question: what kind of theoretical contribution is being offered, and what standards should be used to assess it? Clarifying whether the annulus is intended as a heuristic, a formal model, or a testable theory would significantly strengthen the manuscript.

2. Coherence of the “innovation annulus” as a concept

The annulus concept is the organising device of the paper, but it appears to carry multiple analytical roles simultaneously. It functions as a geometric representation of data availability, a proxy for production inefficiency, a measure of commercial opportunity, and a normative target for governance. It is not clear that a single construct can coherently sustain all of these functions. The boundaries of the annulus (inner and outer radii) are not operationalised in a way that would permit empirical identification, nor is it clear that they are stable across contexts. This raises the possibility that the annulus operates more as a metaphor than a formal analytical model. If so, the limits of that metaphor, and the conditions under which it is informative, should be made explicit.

3. Limited engagement with falsifiability and predictive power

Connected to the above, the framework’s empirical status is underdeveloped. In its current form. It is unclear what empirical observations would disconfirm the model. Many dynamics described (e.g. that more complex data are costlier to produce) risk being tautological. The framework appears able to accommodate a wide range of observed outcomes, raising concerns about explanatory constraint.

The scientific credibility of the argument would be strengthened by more explicitly specifying observable proxies for annulus boundaries, the conditions under which predictions might fail, and potential competing explanations.

4. Risk of post hoc rationalisation

The empirical examples used (e.g. Crossref, Dimensions, GRID/ROR) are informative, but their role in the argument is not entirely clear. As presented, they appear to illustrate the annulus concept rather than test or predict it. This gives rise to a concern that the framework may function primarily as a retrospective rationalisation of historically contingent developments, rather than as a predictive or explanatory model. The manuscript would benefit from clearer articulation of whether and how the annulus framework generates novel expectations about future developments.

5. Under-theorisation of power and strategic behaviour

The analysis foregrounds production costs and demand but pays relatively little attention to strategic behaviour by actors, institutional power asymmetries and the political construction of “frictions”. For example, “production friction” is treated as largely exogenous, whereas it may in part reflect strategic choices, standard-setting processes, or market positioning. The role of incumbents in shaping what is structured, standardised, or withheld is only lightly addressed.

Thus, the framework tends to treat frictions as a natural and persistent feature of the system. However, the extent to which data are structured at source is itself shaped by institutional and economic incentives, and the location and magnitude of “friction” may reflect deliberate design choices rather than inherent constraints. This suggests that the annulus is not simply a response to technical conditions, but also to institutional arrangements and governance decisions, which deserve more explicit treatment.

In short, a more explicit engagement with the political economy of knowledge infrastructures would strengthen the explanatory depth of the model.

6. Narrow (paper-centric) conception of the research system

The framework is strongly oriented around the scholarly record in its conventional (publication-centred) form. While the author notes the historical structure of the record, the analysis does not adequately account for research outputs such as data, software, and protocols. It downplays emerging forms of dissemination and evaluation, and infrastructures not tied to traditional publishing workflows. This raises concerns about the generality of the model. If key components of contemporary research practice are outside the scope of the annulus, its explanatory reach may be significantly limited.

7. Treatment of value in the welfare framework

The welfare-theoretic section introduces value functions (V, B, etc.), but these are not fully specified. It is unclear whose welfare is being modelled. Different types of value (economic, epistemic, social) are implicitly treated as commensurable. The distributional implications of different annulus configurations are not fully explored. Given the centrality of this framework to the paper’s policy claims, a more explicit account of how “value” is defined and aggregated would be desirable.

8. Limited engagement with alternative institutional models

The manuscript critiques a binary opposition between openness and commercial provision, but at points risks reproducing a similar dichotomy. In particular, it gives relatively limited attention to commons-based governance models (Ostrom is mentioned only in passing), cooperatively governed infrastructures, and regionally distinct publication and data ecosystems (e.g. Latin American models, diamond open access). Engagement with these alternatives would both strengthen the argument and test the generality of the annulus framework.

9. Normative stance and positionality

Although the manuscript acknowledges potential conflicts of interest, the framework may implicitly align with particular institutional or commercial perspectives. For example, the notion of an “optimal” annulus risks naturalising a mixed commercial–open equilibrium. Furthermore, empirical support is drawn significantly from systems closely related to the author’s institutional context, and rest partially on assertions about the financial arrangements related to those systems that cannot be tested because the data are not public. This does not invalidate the argument but suggests the importance of broader empirical grounding and engagement with alternative perspectives.

Minor Comments

  • The manuscript occasionally overstates the degree to which research actors (e.g. institutions) operate in strategically rational, data-driven ways; empirical support for these claims would be helpful.
  • Some claims rely heavily on specific case studies; broader empirical evidence would strengthen generalisation.
  • The role of AI as either a centralising force or a potential leveller could be more fully explored.
  • A number of arguments could be expressed more plainly without reliance on geometric metaphor.
  • There are minor typographical errors and inconsistencies in phrasing throughout.

Competing interests: I am associate director of the UK Reproducibility Network, which places emphasis on collaboration and coordination across the research system to address the complicated and entangled challenges involved in improving research rigour and transparency. I have written this review in a personal capacity.

Leave a comment