Published at MetaROR
September 10, 2026
Table of contents
Lost in transition. Quantifying the funding metadata gap in Crossref
1 Dutch Research Council NWO
2 Austrian Science Fund FWF
Originally published on May 8, 2026 at:
Editors
Kathryn Zeiler
Stephen Pinfield
Editorial assessment
by Stephen Pinfield
The purpose of the article is to estimate how much readily available funding information gets registered in Crossref. The authors compare the funding acknowledgements in a sample of articles against the funding metadata registered in Crossref, breaking down the difference by publisher. Both reviewers found the study to be useful and timely and the publisher-level results important. They agree that the results expand our understanding of open science tools. Both reviewers suggest adding background information. For example, more information about what Crossref is and why disclosing funding sources matters would be helpful. Both also suggest clarifying the methods used to produce the results. The reviewers agree that the authors should offer a thicker interpretation of the reported gap. They also suggest expanding the discussion and recommendations to, for example, account for additional actors. The second reviewer suggests ways to enhance the useability of the publicly available data.
Peer review 1
Anonymous reviewer
The manuscript addresses a timely and relevant topic by examining the coverage of funding metadata in Crossref using publication data from two major European funding agencies. The study has the potential to make a valuable contribution to the literature on open research information and bibliometric infrastructures. My comments are primarily intended to strengthen the conceptual framing of the manuscript, improve the transparency and reproducibility of the methodological description, and better articulate the contribution of the findings to the existing literature.
1. Provide more contextual information about Crossref
Some additional information about Crossref, its members, and the reasons for its creation should be provided so that readers can place this research within a broader context. The manuscript focuses very specifically on the coverage of funding metadata, but the broader institutional and infrastructural context of Crossref is largely absent. One limitation of highly technical research is that it can sometimes lose the wider perspective needed to understand the institutional and sociotechnical conditions under which research metadata are produced, curated, standardized, and reused.
2. Explain more clearly why funding metadata matter
The manuscript would benefit from a stronger literature-based justification of the importance of funding metadata for contemporary bibliometrics, research evaluation, and science studies. What kinds of research questions or evaluation problems can funding information help address? Why is improving the completeness of funding metadata important beyond the technical objective of increasing metadata coverage?
Although the manuscript correctly relates its objectives to the principles of the Barcelona Declaration on Open Research Information, it would be useful to explain more explicitly the specific uses and meanings of funding metadata, which constitute the central object of the study. This conceptual framing would help readers better appreciate the broader significance of the research.
3. Clarify the methodological procedure
The methodological section would benefit from a more transparent description of the complete analytical workflow.
The first paragraph describes the administrative origin of the records but not the research data collection procedure itself. It remains unclear from how the two datasets were extracted from their respective sources. If this information is already provided elsewhere, I apologise for overlooking it, but I would nevertheless recommend clarifying this section. Figure 2 clearly illustrates the filtering process but leaves unanswered the question of the initial dataset. Were these records obtained directly from the internal information systems of NWO and FWF? If so, the manuscript should explain what information these systems contain and how the datasets were extracted for research purposes.
A few additional details on how the data were processed for the analyses and the presentation of the results would also improve the reproducibility of the study, specifying the software or computational environment used and the main data processing steps leading to the reported tables and figures.
It would also be useful to explain why Scopus, Dimensions, and HTML scraping were selected as the reference sources for funding acknowledgement texts, how duplicate or conflicting acknowledgements across these sources were managed, and on what basis this combined dataset is considered the reference (ground truth) against which Crossref funding metadata were evaluated.
Finally, the manuscript should specify the computational environment used for the regular expression matching, although they were probably implemented in Google. Furthermore, since the patterns were developed with the assistance of Claude Sonnet 4.5, the manuscript should briefly describe the prompt or prompting strategy used, as well as how the generated patterns were subsequently validated, refined, and verified by the authors before being applied to the corpus.
4. Strengthen the presentation of the results
The authors report that approximately 24% and 16% of the publications lacked explicit acknowledgements to the funding body. While this is an important finding, explanations for these missing acknowledgements should also be considered. For example, differences between publication versions, reporting practices, or the timing at which publications were registered by grantees could also contribute to the observed gap.
The explanation of the agreements between FWF, NWO, and Crossref at the end of page 9 would be more useful if introduced earlier in the manuscript, as it provides important institutional context for understanding the study and interpreting its findings.
More generally, several issues discussed in the results section would benefit from being anticipated in the introduction and more clearly linked to the research questions. Strengthening these connections would improve the overall coherence of the manuscript.
5. Expand the discussion
The discussion relates the findings primarily to the Barcelona Declaration on Open Research Information and to the authors’ previous work. While these comparisons are appropriate, the manuscript would benefit from a broader engagement with the existing literature.
To better demonstrate the contribution of the study, the authors should explain more explicitly how their findings advance the current state of knowledge in relation to previous research on funding metadata, research information infrastructures, and bibliometric data quality. At present, the discussion does not fully clarify what new knowledge this study contributes beyond confirming or extending the authors’ earlier work.
6. Minor observations
The writing could be improved in several places. Because the manuscript builds upon the authors’ previous work and experience, some formulations include occasional sentence fragments and expressions that could be revised for greater precision. For instance, “Either directly from the author when they submit their manuscript in the submission system.”
Finally, the manuscript is somewhat repetitive, particularly in the introductory sections. For example, the estimate that the funding metadata gap in the analysed corpus is approximately 30% is mentioned both in the middle of page 3 and again towards the end of page 4. Reducing this repetition would improve the readability and argumentative progression of the paper.
Peer review 2
This is a valuable and timely study of the completeness of open funding metadata. It provides concrete, policy-relevant results that will be of considerable interest to funders, publishers, scholarly infrastructures, and researchers working with open bibliographic metadata.
The paper investigates the extent to which funding information associated with publications funded by the Dutch Research Council (NWO) and the Austrian Science Fund (FWF) is made available as structured metadata in Crossref. The authors analyze approximately 5,000 publications reported to NWO and 5,000 publications reported to FWF. They combine funding texts obtained from Scopus, Dimensions, and publisher landing pages with funding metadata retrieved through the Crossref API. Their main finding is that, although NWO or FWF can be identified in the funding acknowledgements of approximately 76% and 84% of the publications respectively, the corresponding Funder ID is present in Crossref for only 48% and 55%. The authors therefore identify a funding metadata gap of approximately 30 percentage points. They also document substantial differences between publishers.
The study makes an important contribution to the literature. Existing work, including my own work with Waltman, has documented the share of publications for which publishers make funding information available in Crossref and has revealed large differences between publishers. However, such analyses cannot establish how complete the deposited information is. They do not determine whether publications without Crossref funding metadata nevertheless contain funding information in their full text. Nor do they take into account publications that authors have reported to a funder as resulting from its funding but for which no funding acknowledgement appears in the publication. The present paper addresses these limitations by bringing together information reported by authors to NWO and FWF, funding acknowledgement texts, and Crossref metadata. This combination enables the authors to provide concrete evidence about the extent to which known funding relationships are not represented as structured metadata in Crossref. The study therefore offers a much-needed empirical baseline for assessing and improving the completeness of open funding metadata.
Overall, I consider this a clearly written and useful paper. The research aim is clearly formulated, the methods and data are generally described in a transparent and accessible way, and the results are presented in a manner that is easy to understand. The publisher-level analyses are particularly informative and have clear practical relevance.
I would like to provide the following comments and suggestions for further strengthening this already valuable and important paper.
Definition and interpretation of the metadata gap
The main result is expressed as a metadata gap of about 30 percentage points for both NWO and FWF. This is calculated as the difference between the percentage of all publications in which the funder was found in an acknowledgement and the percentage of all publications for which the corresponding funder ID was found in Crossref. In a sense, however, this measure underestimates the share of available funding information that publishers fail to register in Crossref. Both percentages are calculated relative to the total number of papers reported to the funders, including papers without an identified funding acknowledgement. It could therefore be informative to also report the percentage of publications with an identified acknowledgement of NWO or FWF for which the corresponding funder ID is absent from Crossref. Reporting this conditional measure alongside the current percentage-point gap would provide an additional and intuitive perspective. The current measure shows the size of the metadata gap relative to the full set of reported publications. The conditional measure would show the proportion of publications for which the information is demonstrably present in the acknowledgement but is not represented by the relevant funder ID in Crossref. The terminology used in the paper should make the distinction between these two measures clear.
Regular expressions and validation of funder identification
The analysis depends on regular expressions used to identify NWO and FWF. The regular-expression patterns applied to the collected funding texts are presented in the paper. However, the patterns used to identify mentions of NWO and FWF in the Crossref funder name field are not provided. The paper states only that “a simple regular expression pattern” capturing known variants was used. From the accompanying source code, it appears that different regular expression patterns may have been used for the two sources.
I recommend explaining why different patterns were used for Crossref funder names and full-text funding acknowledgements. If there is no methodological reason for this difference, it may be preferable to use the same patterns. If there is a good reason, the paper should explain the differences.
I also recommend that the authors explain how case sensitivity was handled. The regular expression for NWO presented on page 7 appears to be case-sensitive, whereas the expression for FWF uses the (?i) flag and is therefore case-insensitive. It would be helpful to explain whether this difference was intentional and, if so, why.
Finally, the authors should report whether the regular expressions were tested against a manually coded sample. A modest manual validation exercise that provides an indication of the number of false positives and false negatives, would further increase the robustness and confidence in the findings.
Distinguishing omitted funder information from missing identifiers
Figures 4 and 5 compare mentions of the relevant funder in funding acknowledgements with the presence of the corresponding funder ID in Crossref. This comparison combines two different potential problems: 1) Funding information may be omitted from the Crossref record altogether; 2) A funder name may be deposited, but it may not be normalized to the correct funder ID. The aggregate results already indicate that this distinction matters. I therefore recommend adding, in Figures 4 and 5, the percentage of publications for which funding metadata is available in Crossref and the percentage of publications for which the relevant funder is identified in the Crossref funder name field. This would show whether a publisher’s performance primarily reflects the complete omission of funding information or problems in matching and normalizing a funder name to an identifier.
Furthermore, it would be valuable to analyze whether a funder ID was asserted by the publisher or assigned through Crossref’s matching process. This provenance information is available in the data provided by the Crossref API. The distinction could make the recommendations more actionable. The absence of publisher-asserted identifiers, or a low rate of such identifiers, may indicate that improvements in publisher production workflows are needed. Problems with identifiers assigned through Crossref’s matching process may instead point to opportunities for improving Crossref’s matching and normalization procedures. Such an analysis might also help explain the interesting findings for IOP Publishing and IEEE.
Publisher-level publication counts
The authors indicate that Figures 4 and 5 include the top 20 publishers in terms of publication volume. To facilitate the interpretation of the percentages, it would be helpful to include the absolute number of publications for each publisher.
It would also be useful to indicate what proportion of the complete NWO and FWF datasets is covered by the top 20 publishers. In addition, are there publishers with a significant number of publications that do not deposit relevant funding metadata at all? If so, identifying these publishers, based on a clearly specified minimum number of publications, would provide further insight into the scale and nature of the problem.
Extending the recommendations to funders and DOI registration agencies
The final section offers useful recommendations to publishers and funders. Publishers are encouraged to reassess and improve their practices for submitting funding metadata to Crossref. Funders are encouraged to monitor more systematically whether funded authors comply with grant conditions requiring funding to be acknowledged in publications.
I encourage the authors to extend this discussion. As they describe, funders often maintain detailed information about publications resulting from their funding. The discussion could therefore be broadened beyond improvements to the original publisher deposit workflow.
In particular, the authors could discuss the COMET (Collaborative Metadata) community approach to improving and enriching scholarly metadata, of which I am one of the organizers. The COMET model treats metadata completeness and quality as a shared responsibility. It enables trusted community partners to contribute metadata enrichments together with appropriate provenance information. This approach seems directly relevant to the present study because funders often possess rich internal data about the publications resulting from their funding. Funders such as NWO and FWF could make this data openly available as structured and properly documented metadata enrichments. These enrichments would complement the metadata provided by publishers. Just as publishers are expected to take responsibility for making comprehensive metadata available for the publications they publish, funders could contribute the information they maintain about the outputs resulting from their funding. This would complement, rather than replace, efforts to improve publisher workflows.
An essential element of the COMET model is round-tripping. Community-provided metadata enrichments should flow back into the infrastructures that maintain and disseminate DOI metadata, rather than remaining in disconnected downstream databases. This reduces fragmentation and allows a broad range of downstream systems and users to benefit from the enrichments. COMET and DataCite are already experimenting with such an approach. DataCite maintains enrichment records separately from the metadata provided by the original depositors. These records include provenance information, are subject to validation, and can be used to provide an enriched DOI record without altering the original record.
Another promising recommendation for Crossref and DataCite would be to enable publishers to deposit raw funding acknowledgement text, even when they are unable to transform it into structured funding metadata themselves, as also recommended by Mugabushaka et al. (2022). Publishers may be willing to share funding information but may lack the resources needed to identify funder names accurately and match them to funder IDs. In such cases, making the raw funding text openly available would already be a highly valuable contribution and would be far preferable to losing the information entirely. Crossref, DataCite, or trusted community partners could then take responsibility for extracting funder names, matching them to persistent identifiers, and identifying grant numbers. Retaining the original funding text alongside any derived metadata, with transparent provenance for both, would make it possible to verify and improve these enrichments over time. This would offer a pragmatic and scalable division of responsibilities: publishers make the information they possess available, while shared infrastructures and trusted community partners provide the specialist matching and enrichment capabilities.
I therefore encourage the authors to broaden their discussion and recommendations to address the complementary responsibilities of publishers, funders, DOI registration agencies such as Crossref and DataCite, and trusted community partners. This broader ecosystem perspective would strengthen the policy relevance of the paper. It recognizes that improving funding metadata should be a collective effort and should not depend exclusively on a single actor or a single point in the scholarly communication workflow.
Data and code
It is commendable that the authors have made available on Zenodo the data they are legally permitted to share, together with the source code used for DOI cleaning, Crossref API retrieval, and scraping of publisher landing pages. This is an important strength of the study and demonstrates a strong commitment to openness and transparency. The paper also clearly explains that the funding acknowledgement texts retrieved from Scopus and Dimensions cannot be shared because they come from proprietary sources.
The source code of the scripts has been made available in an RTF document. As a minor recommendation, I suggest making the scripts available as Google Apps Script code files using the .gs (Google Script) extension rather than as formatted text in an RTF file. The authors could additionally make the scripts available in a version-controlled code-sharing environment such as GitHub or GitLab. This would make the code easier to inspect, execute, improve, and reuse.








