Published at MetaROR

September 10, 2026

Table of contents

Cite this article as:

de Jonge, H., & Rieck, K. (2026, May 5). Lost in transition. Quantifying the funding metadata gap in Crossref. Retrieved from osf.io/preprints/metaarxiv/3zm5r_v1

Lost in transition. Quantifying the funding metadata gap in Crossref

Hans de Jonge1EmailORCID, Katharina Rieck2EmailORCID

1 Dutch Research Council NWO
2 Austrian Science Fund FWF

Originally published on May 8, 2026 at: 

Abstract

The importance of having funding metadata openly available is widely acknowledged for research transparency and tracking research funding outcomes. Since 2013, Crossref has provided members the opportunity to deposit funding information when registering DOIs for publications. However, earlier research shows this information is far from complete, with coverage varying significantly across publishers despite funding information often being available in article acknowledgement sections.

This paper quantifies this metadata gap at approximately 30%. Using publications from two national funding councils (the Dutch research council NWO and the Austrian Science Fund FWF), we demonstrate that 30% of funding acknowledgements readily available to publishers in full-text articles are not transferred as metadata to Crossref. We also observe considerable differences between publishers.

This work provides, for the first time, a concrete baseline for improving funding metadata quality and coverage in Crossref. It may also inform publishers about their performance, many of whom may be unaware of these gaps due to outsourced metadata extraction processes.

1. Introduction

The importance of publicly available metadata on the funding of scientific outputs is widely recognized. This is important for transparency and scientific integrity (Nosek et al. 2026; COPE Council, 2025) but from the moment funders started requiring their grantees to acknowledge funding in their manuscripts (late 1990s) this was also done to be able to track funded research outputs (Álvarez-Bornstein & Montesi, 2021). Publication data collected can inform funders’ strategies and are a tangible source to measure the impact of research funding (Álvarez-Bornstein & Montesi, 2021). As funders face growing accountability for the allocation of public resources, many now mandate that grantees formally acknowledge their support in research outputs.

Since 2013, Crossref has offered its members the option of including funding information in the DOI metadata of a publication (Meddings, 2013; Hendricks et al. 2020; Lammey, 2024), see Figure 1. Initially it was recommended to register the name of the organization, the funder_id, and the award. In the beginning of 2026, this schema was extended to include the possibility to register grant DOIs in all types of research output (Feeney, 2026). This has created a potentially very important – open – source for tracking outputs of funded research.

Crossref is therefore regarded as an increasingly interesting source of open bibliographic metadata (Van Eck & Waltman, 2025) and used extensively as a source for downstream bibliographic databases such as Lens, OpenAIRE, Dimensions (Herzog et al., 2020) and OpenAlex (Priem, 2022).

Figure 1. Funding metadata registered in Crossref for Kolmayer et al. (2024) (https://doi.org/10.1038/s41467-024-46937-x). Funding is acknowledged using the funder ID and funder name, identifying the Austrian Science Fund. The award is identified with a grant ID (https://doi.org/10.55776/DOC50).

Registering funding metadata as part of publication metadata in Crossref is not mandatory, and previous research has shown that this information is far from complete (Habermann, 2019; Kramer & De Jonge, 2021; Mugabushaka et al., 2022; Van Eck & Waltman, 2025). Some publishers register funding information for all their publications, but there are also many publishers who only do so for part of their journal portfolio, or not at all. This occurs despite the fact that publishers often have access to this data: namely, as part of the funding acknowledgment section of the full text, where authors disclose their funding in compliance with their funders’ requirements.

The aim of this study is to quantify the gap between the funding information readily available to publishers in the articles they publish and the funding metadata they register in Crossref.

We use two large publication datasets from two national research funding councils: the Dutch Research Council (NWO) and the Austrian Science Fund (FWF). We estimate the discrepancy to be around 30%. In other words, for 30% of the publications in this dataset, funding metadata is missing in Crossref, even though the articles themselves contain a very clear reference to the funding organization. Somehow, however, this information does not reach Crossref. We observe interesting differences between publishers.

The importance of this research lies in the growing interest in open bibliographic metadata, sparked by initiatives such as the Barcelona Declaration on Open Research Information (2024). The aim of this initiative is to promote the availability of open research information, especially open bibliographic metadata. The argument for this is that institutions, researchers and society at large, should not be dependent on closed, often selective, biased, and proprietary databases.

Reducing reliance on these databases requires inter alia improving the quality and completeness of funding information in open infrastructures such as Crossref. Increasingly, publisher negotiations are seen as an important mechanism for achieving this improvement (Barcelona Declaration, 2025), but for that, funders and libraries need clear benchmarks. This paper provides such a baseline for the first time. This work may also inform publishers about their performance in registering funding metadata with Crossref, many of whom may be unaware of these gaps.

2. Earlier research

While the literature on funding acknowledgements, their use and limitations in policy analysis and bibliometric research is vast (Álvarez-Bornstein & Montesi, 2021), studies on the completeness of funding metadata in Crossref are limited. Van Eck & Waltman (2025) have tracked the progress of six different metadata types in Crossref, including funding metadata, since 2021. They have shown that the percentage of publications in Crossref that include some form of funding metadata has increased to 25% in 2024.1 That percentage, however, was already reached in 2020 and has not increased since. They also show that there is substantial variation between publishers, with some larger society publishers including funding metadata for almost all of their publications, while others – including some of the larger commercial publishers – share this information only for a subset of their journals or not at all.

Similar conclusions were reached by Mugabashaka et al. (2022) who compared the coverage of funding metadata in Crossref with both Web of Science and Scopus for a large corpus of publications on Covid-19.

The present study builds on a paper by Kramer & De Jonge (2022). The aim of that study was to determine the quality and completeness of funding metadata in Crossref. Based on a corpus of around 5,000 NWO-funded publications, it concluded that just over half referenced NWO by name in Crossref funding metadata and just under half included the funder ID. Comparing this with other bibliographic databases (Web of Science, Scopus, Dimensions and Lens) showed that these commercial databases were able to infer funding information for a substantial number of additional publications that lacked funding metadata in Crossref. It was concluded this extra data was extracted from the acknowledgement section of publications by means of text and data mining.

In a workshop organised as a follow-up to this work together with Crossref, publishers confirmed that they are often struggling with the collection and retention of funding metadata through their production workflows (De Jonge et al. 2023). Broadly speaking, there are two ways in which publishers collect this information. Either directly from the author when they submit their manuscript in the submission system. Other publishers rely on extracting funding metadata from the funding acknowledgement sections of the article further down the production line. They often outsource this to third-party vendors and leave the extraction and submission of funding metadata to Crossref to these parties – a practice which Crossref allows in principle (Crossref, n.d.).

Publishers with exceptionally complete funding metadata in Crossref operate workflows that combine both approaches. Funding information is collected from authors at submission and subsequently compared with information from the manuscript. When differences or conflicts are being noted, the data is fixed, completed, and reformatted and fed back to the author for validation (De Jonge et al. 2023).

In this paper, our focus is on the discrepancy between the funding information available to publishers in the funding acknowledgements section of publications and the funding metadata available in Crossref. Using two large sets of publication data from the Dutch Research Council NWO and the Austrian Science Fund (FWF) we quantify this funding metadata gap to be almost 30%. In other words: for nearly a third of publications of both NWO and FWF funding information is readily available in the funding acknowledgement section of the full text – NWO and FWF are both clearly identified as the funders – but somehow this information is not transferred to Crossref or is lost in that transition.

3. Method

3.1. Data collection

For this analysis we made use of two data sets containing all peer reviewed articles reported by grantees to the Dutch Research Council NWO and the Austrian Science Fund FWF in one year as part of their end-of-grant reports. Both NWO and FWF require grantees to include funding acknowledgements in all research outputs and to report these publications in dedicated systems. NWO uses its grant management system ISAAC, FWF makes use of ResearchFish, an Elsevier-owned platform specifically developed to allow funding organisations to track funded output (Hinrich et al. 2015). For reasons of availability, for NWO all journal articles registered in 2025 were used, while for FWF we used publications reported in 2024.

Because grant holders do not necessarily register their publications in the year of publication, both datasets contain articles published in multiple years. For NWO this spans the years 2013 to 2026, for FWF the years 2014 to 2025.

Cleaning and deduplication of the data followed the same procedure for both funders. Journal articles without a DOI were left out of the analysis. DOIs were cleaned using a Google script (De Jonge, 2026), stripping them from url-prefixes, trailing punctuation marks and other invalid characters. As publications can be reported as outputs of multiple projects, data was deduplicated. Finally, we excluded publications with DOIs from agencies other than Crossref (such as DataCite), as our analysis focuses specifically on Crossref metadata. This resulted in a final dataset of 5,048 publications for NWO and 5,312 for FWF. Figure 2 below shows how many records were excluded at each stage in the process.

Figure 2. Composition of the dataset for NWO- and FWF-funded DOIs used in this analysis

3.2. Retrieval of funding metadata from Crossref

For all records in the data set, metadata including funding information was retrieved using the Crossref REST API. The API was queried through a Google Apps script (De Jonge, 2026), returning the results directly to Google Sheets for further processing. Metadata retrieved for each publication included: member id, publication year and the publisher name.

Subsequently, for each publication, the following funding metadata was retrieved:

  • Whether the record included any funding information at all
  • The funder_name
  • The funder_id
  • The award

For the funder name extraction a simple regular expression pattern was used that captures most known name variants of FWF and NWO in English, Dutch and German. Records were queried for the presence or absence of the following funder IDs associated with NWO and FWF:

  • 13039/501100003246 (Nederlandse Organisatie voor Wetenschappelijk Onderzoek)2
  • 13039/501100024871 (Sociale en Geesteswetenschappen, NWO)
  • 13039/501100024872 (Toegepaste en Technische Wetenschappen, NWO)
  • 13039/501100024870 (Exacte en Natuurwetenschappen)
  • 13039/501100010071 (Nationaal Regieorgaan Onderwijsonderzoek)
  • 13039/501100010409 (Nationaal Regieorgaan Praktijkgericht Onderzoek SIA)
  • 13039/501100002428 (FWF Austrian Science Fund)3

3.3. Retrieval of funding texts from full text

For all records in both datasets an attempt was made to retrieve the funding acknowledgement text as it is included in the full text. In order to obtain the largest possible number of funding texts a combination of three sources were used. Using NWO’s licensed instance of Scopus (Baas et al. 2020) we collected the funding acknowledgements for as many records as possible. Raw funding acknowledgement strings are provided by Scopus in the field “Funding Text”. In addition funding information was retrieved for as many records as possible using FWF’s licensed instance of Dimensions (Herzog et al. 2020). We retrieved the fields “Funding” and “Acknowledgement” as these both contain funding information. In addition, we used a Google script (De Jonge, 2026), to scrape the HTML landing pages of articles. This yielded a small number of additional funding declarations for which no funding information could be retrieved either with Scopus or Dimensions (69 for NWO,18 for FWF).

To identify mentions of NWO and FWF in funding texts, we used regular expression patterns developed with the assistance of Claude Sonnet 4.5 (Anthropic, 2025). The patterns were designed to capture known name variants in English, Dutch, and German, as well as program-specific terminology (e.g., VENI, VIDI, VICI for NWO; START, Wittgenstein for FWF). The patterns were:

  • REGEXMATCH(H2; “N\.?W\.?O\.?|Netherlands Organi|Nederlandse Organis|Dutch Research Council|STW|VENI|VIDI|VICI|ALW|FOM|TTW|Spinoza|Gravitation|Aard-en Levenswetenschappen|Fundamenteel Onderzoek”)
  • REGEXMATCH(H2; “(?i)f\.?w\.?f\.?|austrian science fund|austrian (science|research) (fund|foundation|council)|austrian (national science fund|foundation for scientific research)|fonds zur f(ö|oe|o)rderung|(ö|oe)sterreichische.*wissenschaft| wissenschaftsfonds|austrianresearchpromotion|\bstart.?programm(e)?\b|wittgenstein.?( preis|prize|award)|erwin.?schr(ö|oe)dinger|lise.?meitner|elise?\.?richter|hertha.?firnber g|\bdoc\b|\besp\b.*programm”)

4. Results

4.1. The gap between funding acknowledgement and funding metadata in Crossref

It is worth recalling that all publications in this study have been reported by researchers as being the result of funding they received from NWO or FWF. Both funding bodies require their grantees to acknowledge this funding in their project output including publications.4 In principle, therefore, 100% of these articles should contain a reference to NWO or FWF in the funding acknowledgement section. However, Figure 3 shows that researchers often fail to do so. Acknowledgement of NWO is missing in the acknowledgement section of 23.9% of publications, while FWF is missing in 15.9% of publications.

This means that NWO and FWF are mentioned by name in the funding acknowledgement in 76.0% and 84.1% of publications respectively. However, when we look at how much of this information is submitted by publishers as structured open funding metadata to Crossref, the gap is substantial. For NWO, only 48,3% of the publication’s metadata contain an NWO funder ID. For FWF, this is 54,8%. This represents a funding metadata gap of 27.8 percentage points for NWO and 29.3 percentage points for FWF. In other words, for almost 30% of the publications in this dataset, funding information was available – unambiguously identifying NWO or FWF by name in the acknowledgement section – yet this information was not consistently deposited as structured funding metadata to Crossref. This is not a matter of missing or ambiguous information on the author’s side, it suggests a structural issue in the publisher metadata pipeline.

The situation is slightly better when we look at the extent to which NWO and FWF are identified in Crossref with the name of their organisation. The metadata discrepancy between acknowledgement section and funding metadata in Crossref is then 22.9% (NWO) and 21.2% (FWF) respectively. Given the known variation in organisation names also for funders like NWO and FWF, it is of course highly preferable to identify funders with an organisational ID (Funder ID or ROR ID).

When we compare these results with an earlier analysis (Kramer & De Jonge, 2022) from almost five years ago, we must also conclude that not much progress has been made in the meantime. Out of the 5,004 NWO funded publications analysed in 2022 only 45% were correctly attributed to NWO with the use of its funder ID (compared to 48.3% for the current dataset).

Figure 3. The discrepancy between available funding information in full text and Crossref for NWO and FWF. The first two bars indicate whether a funding acknowledgement was retrieved and whether NWO or FWF was mentioned in it. The following bars indicate the extent to which Crossref has (any) funder metadata, the extent to which NWO or FWF were identified by their organisation name, their funder ID, and whether the funding metadata contained any award data.

4.2. Funding awards and grant ID’s

Crossref allows its members to register grant identifiers as part of the publication metadata. The benefit is that publications can be linked to specific awards. The number of publications that include metadata about awards is comparable to the amount of funder IDs. For NWO, 43% of publications contain award metadata, whereas for FWF this figure is 55.1%.

Both FWF and NWO joined Crossref’s Grant Linking System (GLS) and assign DOIs to awarded projects so that their funding can be uniquely identified globally (Crossref, n.d.). FWF introduced grant DOIs in 2023 and has done so for all FWF projects, including those approved as far back as 1995, whereas NWO decided to do so for 2024 onwards only (FWF, 2023, NWO, 2024). Since then, both funders require authors to acknowledge their funded projects in research outputs using the grant DOI rather than their own award numbers / IDs. The available data on the use of grant DOIs to acknowledge research funding is therefore still very limited and not representative. The extent to which the grant DOIs are included in the metadata of articles is still very limited. For NWO, we found only three grant DOIs in acknowledgement sections in this dataset, one of which is translated in Crossref metadata. For FWF, we found 210 grant DOIs in acknowledgement sections, 123 are also included in the Crossref metadata. The overall low number of grant IDs being submitted to Crossref is consistent with earlier observations (De Jonge, 2025).

4.3. Publisher performance for NWO funded papers

We observe considerable differences between publishers. Figure 4 shows the difference between the frequency with which NWO is mentioned in funding acknowledgements and the frequency with which that information also ends up as structured metadata – a funder ID – in Crossref for the top 20 publishers in terms of publication volume.

In line with previous analyses (Kramer & De Jonge, 2022, Van Eck & Waltman, 2025), three major society publishers stand out for their strong performance: the American Chemical Society (ACS), the Royal Chemical Society (RSC) and the American Institute of Physics (AIP) all show very small discrepancies. In almost 90% of their articles, NWO is correctly cited as the funder, and in almost 80% of cases, NWO is also identified in Crossref with a funder ID. These publishers appear to have mature, well-integrated workflows that reliably capture and transfer funding information to Crossref.

However, the picture changes significantly for other major publishers. At Springer Nature the gap is 39.4%, at Oxford University Press 34.3%, and at MDPI 43.9%, meaning that for roughly one in three publications from these publishers, funding information present in the article does not reach Crossref. The situation is more concerning still for EDP Sciences and IEEE where the discrepancy reaches as high as 60%. For these publishers, the majority of available funding information is not included in the Crossref metadata- despite being present in the acknowledgement sections of the articles themselves.

Figure 4. The discrepancy between funding information available in full text and funding metadata available in Crossref split out by publisher for NWO funded publications.

4.4. Publisher performance for FWF funded papers

Figure 5 shows, for the top 20 publishers by publication volume, the discrepancy between how often FWF is mentioned in funding acknowledgements and how frequently this information is also deposited as structured metadata with a funder ID in Crossref . The pattern closely mirrors what was observed for NWO-funded publications.

The society publishers American Physical Society (APS), American Chemical Society (ACS), American Institute of Physics (AIP), and Copernicus all show very low gap scores, indicating that they consistently transfer funding information available in acknowledgement texts into Crossref metadata. For Elsevier, Oxford University Press, PLOS and Royal Society of Chemistry the gap is around 20%, indicating room for improvement. For MDPI, Wiley,

Springer Nature and Frontiers, the gap widens considerably to over 30%, meaning that for more than one in three FWF-funded publications, funding information present in the article is not included in Crossref metadata.

Most striking are the results for EDP Sciences, Cambridge University Press and IOP Publishing, where the gap reaches between 60% and 75%. For these publishers, the vast majority of available funding information is lost in the transition from full text to metadata.

Figure 5. The discrepancy between funding information available in full text and funding metadata available in Crossref split out by publisher for FWF.

4.5. Metadata gaps compared between NWO and FWF

In Figure 6 we compare the metadata gaps between acknowledgement section and Crossref metadata for the 18 largest publishers between NWO and FWF. In general, we see a very similar pattern across publishers. EDP Sciences and Cambridge University Press each show a gap of approximately 50% for both funders and with that the highest difference, whereas the American Physical Society (APS), the American Chemical Society (ACS), the Royal Society of Chemistry (RSC) and Copernicus perform very well with very small discrepancies between the funding information shared by authors in acknowledgements and what is transferred to Crossref.

The only notable outliers are IOP and IEEE. IOP appears to have difficulty translating funding information into the FWF funder ID, while IEEE struggles to identify NWO. Both publishers are able to register the funder’s organisation name but fail to register the corresponding funder ID. We examined whether journal-level variation or differences in publishing behaviour between NWO- and FWF-funded researchers might explain this pattern, but found no evidence for either.

Figure 6. Comparing the metadata gap between funding acknowledgement and presence of funding metadata in Crossref for NWO and FWF per publisher.

5. Conclusion and policy recommendations

Previous research has already shown that the funding metadata in Crossref is incomplete and that there are major differences between publishers.

The aim of this study was to quantify the gap between funding information that is readily available in acknowledgement sections of research articles but nevertheless does not find its way into Crossref as structured data. Our analysis shows that this gap is approximately 30% – meaning that for nearly a third of publications, NWO and FWF were clearly acknowledged as funders but this was not reflected in Crossref as structured metadata. We observe important differences between publishers, with some sharing up to 90% of funding information included in the acknowledgement sections of articles with Crossref, whereas others perform considerably less well. Future research should determine whether these findings generalize to other funding organizations.

The significance of these findings lies in the growing importance of open bibliographic metadata, as evidenced by the rapid adoption of the Barcelona Declaration on Open Research Information. In this context, the importance of open funding metadata is widely recognized (Barcelona Declaration, 2024). It ensures transparency regarding how and by which organization research is funded. Open funding metadata is also relevant for tracking the outputs from the funding provided by funding bodies such as NWO and FWF. This information is important as a source for strategy building, but also as a source for accountability for publicly funded research. Therefore, to be truly useful, completeness of open funding metadata in databases such as Crossref is of great importance.

We hope that the results will encourage publishers to critically reassess their current practices regarding the submission of funding metadata to Crossref. As De Jonge & Kramer (2026) argued elsewhere, technical capabilities of publishers play an important role in their ability to register metadata to Crossref, which places considerable demands on the interoperability of the submission and production systems they use. Other publishers are known to have outsourced the extraction and submission to Crossref of funding metadata to third-party vendors They would do well to more systematically audit the performance of these vendors.

We also hope that our research helps inform research institutions that negotiate publishing agreements in establishing benchmarks for funding metadata as part of open bibliographic metadata.These negotiations are increasingly seen as an important means of improving the availability of bibliographic metadata (Barcelona Declaration, 2025).

Our findings also show that funders themselves could do more. Our research shows that, despite the fact that the acknowledgement of funding is a grant condition of both funders, a substantial number of publications in our dataset did not contain a reference to the funding of NWO or FWF.

Our research shows that 30% of funding information is lost in transition. The infrastructure to address this issue is already in place, and this article provides clear benchmark figures for making improvements. We hope that this research will help bridge this gap and, in doing so, improve the quality of funding metadata.

Acknowledgements

The authors used Deepl to improve the readability and language of this manuscript. The authors reviewed and edited the output and take full responsibility for the content of this publication.

Code development was assisted by Claude Sonnet 4.5 (Anthropic, 2025). The authors reviewed, tested, and validated all code to ensure correctness and appropriateness for the research objectives.

Author contributions

CRediT: Conceptualization: HDJ; Data curation: HDJ; Formal Analysis: HDJ; Funding acquisition: ; Investigation: HDJ, KR; Methodology: HDJ, KR; Project administration: HDJ; Resources: ; Software: HDJ; Supervision: HDJ; Validation: HDJ, KR; Visualization: HDJ; Writing – original draft: HDJ; Writing – review & editing: KR

Competing interests 

HdJ and KR are co-chairs of the Funding Metadata Working Group under the Barcelona Declaration of Open Research Information. KR is an elected member of the Crossref board.

The authors write in a personal capacity and views they share in this article do not necessarily express the opinion of their employers.

Funding information 

The authors did not receive any funding for the research reported in this paper.

Data availability

Data and code are openly shared to the extent possible. Funding acknowledgement texts retrieved from Scopus and Dimensions cannot be shared, as these are proprietary systems. However, the derived indicators – whether a funding acknowledgement text was found in either system (true/false) – are included. The following information is available via Zenodo:

  • Publication records for NWO and FWF including associated Crossref
  • Google apps scripts to clean DOIs, retrieve Crossref funding metadata and scrape HTML landing pages in search for funding acknowledgements.

Notes

1 An interactive version of the figures they present is available at: https://tinyurl.com/3zk5nvvf.
2
See: https://api.crossref.org/funders/501100003246
3 See: https://api.crossref.org/funders/501100002428
4 For NWO see: https://www.nwo.nl/en/acknowledgement-in-publications, for FWF see: https://www.fwf.ac.at/en/about-us/what-we-do/open-science/open-access-policy/open-access-policy-for-peer-reviewed-book-publications

References

Álvarez-Bornstein, B., & Montesi, M. (2021). Funding acknowledgements in scientific publications: A literature review. Research Evaluation, 29(4), 469–488. https://doi.org/10.1093/reseval/rvaa038

Anthropic. (2025). Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5

Baas, J., Schotten, M., Plume, A., Côté, G., & Karimi, R. (2020). Scopus as a curated, high-quality bibliometric data source for academic research in quantitative science studies. Quantitative Science Studies, 1(1), 377–386. https://doi.org/10.1162/qss_a_00019

Barcelona Declaration on Open Research Information. (2024). Barcelona Declaration on Open Research Information. https://doi.org/10.5281/ZENODO.10958521

Barcelona Declaration on Open Research Information. (2024). Report of the Paris Conference on Open Research Information. https://doi.org/10.5281/zenodo.14054243

Barcelona Declaration on Open Research Information. (2025). Barcelona Declaration and OA2020 launch joint task force on negotiating openness of publication metadata. https://barcelona-declaration.org/news/20251002_bd_oa2020_joint_taskforce/

COPE Council. (2025). COPE discussion document: Declaring funding sources for research. https://doi.org/10.24318/OUWWQgSb

Crossref. (n.d.). Grant linking system (GLS). https://www.crossref.org/services/grant-linking-system/

Crossref. (n.d.). Working with a service provider. https://www.crossref.org/documentation/member-setup/working-with-a-service-provider/

Van Eck, N. J., & Waltman, L. (2025). Crossref as a source of open bibliographic metadata. https://doi.org/10.31222/osf.io/smxe5_v2

Feeney, P. (2026). The best way of acknowledging research funding in the metadata: Crossref Grant ID. https://doi.org/10.64000/x7d4h-x3r11

FWF. (2023). New identification numbers for FWF projects. https://www.fwf.ac.at/en/news/detail/neue-identifikations-nummer-fuer-fwf-projekte

Habermann, T. (2019). The big picture — Has CrossRef metadata completeness improved? Metadata Game Changers. https://metadatagamechangers.com/blog/201

Hendricks, G., Tkaczyk, D., Lin, J., & Feeney, P. (2020). Crossref: The sustainable source of community-owned scholarly metadata. Quantitative Science Studies, 1(1), 414–427. https://doi.org/10.1162/qss_a_00022

Herzog, C., Hook, D., & Konkiel, S. (2020). Dimensions: Bringing down barriers between scientometricians and data. Quantitative Science Studies, 1(1), 387–395. https://doi.org/10.1162/qss_a_00020

Hinrichs, S., Montague, E., & Grant, J. (2015). Researchfish: A forward look. Challenges and opportunities for using Researchfish to support research assessment. The Policy Institute at King’s College London. https://s3.amazonaws.com/rf-downloads/Kings+College+Report.pdf

De Jonge, H., Kramer, B., Michaud, F., & Hendricks, G. (2023). Open funding metadata through Crossref: A workshop to discuss challenges and improving workflows. Front Matter. https://doi.org/10.64000/3f63f-yt393

De Jonge, H. (2025). Celebrating one year of Crossref grant IDs at NWO. Front Matter. https://doi.org/10.64000/dvqke-j4v69

De Jonge, H., & Kramer, B. (2026). Manuscript submission systems and metadata completeness in Crossref: Patterns and associations. PLOS ONE, 21(3), e0345417. https://doi.org/10.1371/journal.pone.0345417

De Jonge, H. (2026). Dataset: Lost in transition. Quantifying the funding metadata gap in Crossref. https://doi.org/10.5281/zenodo.19475976

Kohlmayr, J. M., Grabner, G. F., Nusser, A. (2024). Mutational scanning pinpoints distinct binding sites of key ATGL regulators in lipolysis. Nature Communications, 15, Article 2516. https://doi.org/10.1038/s41467-024-46937-x

Kramer, B., & de Jonge, H. (2022). The availability and completeness of open funder metadata: Case study for publications funded by the Dutch Research Council. Quantitative Science Studies, 3(2), 1–24. https://doi.org/10.1162/qss_a_00210

Lammey, R. (2014). CrossRef developments and initiatives: An update on services for the scholarly publishing community from CrossRef. Science Editing, 1(1), 13–18. https://doi.org/10.6087/kcse.2014.1.13

Meddings, K. (2013). FundRef: Connecting research funding to published outcomes. Insights, 26(3), 272–276. https://doi.org/10.1629/2048-7754.98

Mugabushaka, A.-M., Van Eck, N. J., & Waltman, L. (2022). Funding COVID-19 research: Insights from an exploratory analysis using open data infrastructures. Quantitative Science Studies, 3(3), 560–582. https://doi.org/10.1162/qss_a_00212

Nosek, B. A., Allison, D. B., Jamieson, K. H., McNutt, M., Nielsen, A. B., & Wolf, S. M. (2026). A framework for assessing the trustworthiness of scientific research findings.

Proceedings of the National Academy of Sciences, 123(6), e2536736123. https://doi.org/10.1073/pnas.2536736123

NWO. (2024). NWO funded research projects get unique identifier with grant ID. https://www.nwo.nl/en/news/nwo-funded-research-projects-get-unique-identifier-with-grant-i d

Priem, J., Piwowar, H., & Orr, R. (2022). OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. https://doi.org/10.48550/arXiv.2205.01833

Editors

Kathryn Zeiler
Editor-in-Chief

Stephen Pinfield
Handling Editor

Editorial assessment

by Stephen Pinfield

DOI: 10.70744/MetaROR.431.1.ea

The purpose of the article is to estimate how much readily available funding information gets registered in Crossref. The authors compare the funding acknowledgements in a sample of articles against the funding metadata registered in Crossref, breaking down the difference by publisher. Both reviewers found the study to be useful and timely and the publisher-level results important. They agree that the results expand our understanding of open science tools. Both reviewers suggest adding background information. For example, more information about what Crossref is and why disclosing funding sources matters would be helpful. Both also suggest clarifying the methods used to produce the results. The reviewers agree that the authors should offer a thicker interpretation of the reported gap. They also suggest expanding the discussion and recommendations to, for example, account for additional actors. The second reviewer suggests ways to enhance the useability of the publicly available data.

Competing interests: None.

Peer review 1

Anonymous reviewer

DOI: 10.70744/MetaROR.431.1.rv1

The manuscript addresses a timely and relevant topic by examining the coverage of funding metadata in Crossref using publication data from two major European funding agencies. The study has the potential to make a valuable contribution to the literature on open research information and bibliometric infrastructures. My comments are primarily intended to strengthen the conceptual framing of the manuscript, improve the transparency and reproducibility of the methodological description, and better articulate the contribution of the findings to the existing literature. 

1. Provide more contextual information about Crossref 

Some additional information about Crossref, its members, and the reasons for its creation should be provided so that readers can place this research within a broader context. The manuscript focuses very specifically on the coverage of funding metadata, but the broader institutional and infrastructural context of Crossref is largely absent. One limitation of highly technical research is that it can sometimes lose the wider perspective needed to understand the institutional and sociotechnical conditions under which research metadata are produced, curated, standardized, and reused.  

2. Explain more clearly why funding metadata matter 

The manuscript would benefit from a stronger literature-based justification of the importance of funding metadata for contemporary bibliometrics, research evaluation, and science studies. What kinds of research questions or evaluation problems can funding information help address? Why is improving the completeness of funding metadata important beyond the technical objective of increasing metadata coverage? 

Although the manuscript correctly relates its objectives to the principles of the Barcelona Declaration on Open Research Information, it would be useful to explain more explicitly the specific uses and meanings of funding metadata, which constitute the central object of the study. This conceptual framing would help readers better appreciate the broader significance of the research. 

3. Clarify the methodological procedure 

The methodological section would benefit from a more transparent description of the complete analytical workflow. 

The first paragraph describes the administrative origin of the records but not the research data collection procedure itself. It remains unclear from how the two datasets were extracted from their respective sources. If this information is already provided elsewhere, I apologise for overlooking it, but I would nevertheless recommend clarifying this section. Figure 2 clearly illustrates the filtering process but leaves unanswered the question of the initial dataset. Were these records obtained directly from the internal information systems of NWO and FWF? If so, the manuscript should explain what information these systems contain and how the datasets were extracted for research purposes. 

A few additional details on how the data were processed for the analyses and the presentation of the results would also improve the reproducibility of the study, specifying the software or computational environment used and the main data processing steps leading to the reported tables and figures. 

It would also be useful to explain why Scopus, Dimensions, and HTML scraping were selected as the reference sources for funding acknowledgement texts, how duplicate or conflicting acknowledgements across these sources were managed, and on what basis this combined dataset is considered the reference (ground truth) against which Crossref funding metadata were evaluated. 

Finally, the manuscript should specify the computational environment used for the regular expression matching, although they were probably implemented in Google. Furthermore, since the patterns were developed with the assistance of Claude Sonnet 4.5, the manuscript should briefly describe the prompt or prompting strategy used, as well as how the generated patterns were subsequently validated, refined, and verified by the authors before being applied to the corpus. 

4. Strengthen the presentation of the results 

The authors report that approximately 24% and 16% of the publications lacked explicit acknowledgements to the funding body. While this is an important finding, explanations for these missing acknowledgements should also be considered. For example, differences between publication versions, reporting practices, or the timing at which publications were registered by grantees could also contribute to the observed gap. 

The explanation of the agreements between FWF, NWO, and Crossref at the end of page 9 would be more useful if introduced earlier in the manuscript, as it provides important institutional context for understanding the study and interpreting its findings. 

More generally, several issues discussed in the results section would benefit from being anticipated in the introduction and more clearly linked to the research questions. Strengthening these connections would improve the overall coherence of the manuscript. 

5. Expand the discussion 

The discussion relates the findings primarily to the Barcelona Declaration on Open Research Information and to the authors’ previous work. While these comparisons are appropriate, the manuscript would benefit from a broader engagement with the existing literature. 

To better demonstrate the contribution of the study, the authors should explain more explicitly how their findings advance the current state of knowledge in relation to previous research on funding metadata, research information infrastructures, and bibliometric data quality. At present, the discussion does not fully clarify what new knowledge this study contributes beyond confirming or extending the authors’ earlier work. 

6. Minor observations 

The writing could be improved in several places. Because the manuscript builds upon the authors’ previous work and experience, some formulations include occasional sentence fragments and expressions that could be revised for greater precision. For instance, “Either directly from the author when they submit their manuscript in the submission system.” 

Finally, the manuscript is somewhat repetitive, particularly in the introductory sections. For example, the estimate that the funding metadata gap in the analysed corpus is approximately 30% is mentioned both in the middle of page 3 and again towards the end of page 4. Reducing this repetition would improve the readability and argumentative progression of the paper. 

Competing interests: None.

Peer review 2

Nees Jan van Eck

DOI: 10.70744/MetaROR.431.1.rv2

This is a valuable and timely study of the completeness of open funding metadata. It provides concrete, policy-relevant results that will be of considerable interest to funders, publishers, scholarly infrastructures, and researchers working with open bibliographic metadata.

The paper investigates the extent to which funding information associated with publications funded by the Dutch Research Council (NWO) and the Austrian Science Fund (FWF) is made available as structured metadata in Crossref. The authors analyze approximately 5,000 publications reported to NWO and 5,000 publications reported to FWF. They combine funding texts obtained from Scopus, Dimensions, and publisher landing pages with funding metadata retrieved through the Crossref API. Their main finding is that, although NWO or FWF can be identified in the funding acknowledgements of approximately 76% and 84% of the publications respectively, the corresponding Funder ID is present in Crossref for only 48% and 55%. The authors therefore identify a funding metadata gap of approximately 30 percentage points. They also document substantial differences between publishers.

The study makes an important contribution to the literature. Existing work, including my own work with Waltman, has documented the share of publications for which publishers make funding information available in Crossref and has revealed large differences between publishers. However, such analyses cannot establish how complete the deposited information is. They do not determine whether publications without Crossref funding metadata nevertheless contain funding information in their full text. Nor do they take into account publications that authors have reported to a funder as resulting from its funding but for which no funding acknowledgement appears in the publication. The present paper addresses these limitations by bringing together information reported by authors to NWO and FWF, funding acknowledgement texts, and Crossref metadata. This combination enables the authors to provide concrete evidence about the extent to which known funding relationships are not represented as structured metadata in Crossref. The study therefore offers a much-needed empirical baseline for assessing and improving the completeness of open funding metadata.

Overall, I consider this a clearly written and useful paper. The research aim is clearly formulated, the methods and data are generally described in a transparent and accessible way, and the results are presented in a manner that is easy to understand. The publisher-level analyses are particularly informative and have clear practical relevance.

I would like to provide the following comments and suggestions for further strengthening this already valuable and important paper.

Definition and interpretation of the metadata gap

The main result is expressed as a metadata gap of about 30 percentage points for both NWO and FWF. This is calculated as the difference between the percentage of all publications in which the funder was found in an acknowledgement and the percentage of all publications for which the corresponding funder ID was found in Crossref. In a sense, however, this measure underestimates the share of available funding information that publishers fail to register in Crossref. Both percentages are calculated relative to the total number of papers reported to the funders, including papers without an identified funding acknowledgement. It could therefore be informative to also report the percentage of publications with an identified acknowledgement of NWO or FWF for which the corresponding funder ID is absent from Crossref. Reporting this conditional measure alongside the current percentage-point gap would provide an additional and intuitive perspective. The current measure shows the size of the metadata gap relative to the full set of reported publications. The conditional measure would show the proportion of publications for which the information is demonstrably present in the acknowledgement but is not represented by the relevant funder ID in Crossref. The terminology used in the paper should make the distinction between these two measures clear.

Regular expressions and validation of funder identification

The analysis depends on regular expressions used to identify NWO and FWF. The regular-expression patterns applied to the collected funding texts are presented in the paper. However, the patterns used to identify mentions of NWO and FWF in the Crossref funder name field are not provided. The paper states only that “a simple regular expression pattern” capturing known variants was used. From the accompanying source code, it appears that different regular expression patterns may have been used for the two sources.

I recommend explaining why different patterns were used for Crossref funder names and full-text funding acknowledgements. If there is no methodological reason for this difference, it may be preferable to use the same patterns. If there is a good reason, the paper should explain the differences.

I also recommend that the authors explain how case sensitivity was handled. The regular expression for NWO presented on page 7 appears to be case-sensitive, whereas the expression for FWF uses the (?i) flag and is therefore case-insensitive. It would be helpful to explain whether this difference was intentional and, if so, why.

Finally, the authors should report whether the regular expressions were tested against a manually coded sample. A modest manual validation exercise that provides an indication of the number of false positives and false negatives, would further increase the robustness and confidence in the findings.

Distinguishing omitted funder information from missing identifiers

Figures 4 and 5 compare mentions of the relevant funder in funding acknowledgements with the presence of the corresponding funder ID in Crossref. This comparison combines two different potential problems: 1) Funding information may be omitted from the Crossref record altogether; 2) A funder name may be deposited, but it may not be normalized to the correct funder ID. The aggregate results already indicate that this distinction matters. I therefore recommend adding, in Figures 4 and 5, the percentage of publications for which funding metadata is available in Crossref and the percentage of publications for which the relevant funder is identified in the Crossref funder name field. This would show whether a publisher’s performance primarily reflects the complete omission of funding information or problems in matching and normalizing a funder name to an identifier.

Furthermore, it would be valuable to analyze whether a funder ID was asserted by the publisher or assigned through Crossref’s matching process. This provenance information is available in the data provided by the Crossref API. The distinction could make the recommendations more actionable. The absence of publisher-asserted identifiers, or a low rate of such identifiers, may indicate that improvements in publisher production workflows are needed. Problems with identifiers assigned through Crossref’s matching process may instead point to opportunities for improving Crossref’s matching and normalization procedures. Such an analysis might also help explain the interesting findings for IOP Publishing and IEEE.

Publisher-level publication counts

The authors indicate that Figures 4 and 5 include the top 20 publishers in terms of publication volume. To facilitate the interpretation of the percentages, it would be helpful to include the absolute number of publications for each publisher.

It would also be useful to indicate what proportion of the complete NWO and FWF datasets is covered by the top 20 publishers. In addition, are there publishers with a significant number of publications that do not deposit relevant funding metadata at all? If so, identifying these publishers, based on a clearly specified minimum number of publications, would provide further insight into the scale and nature of the problem.

Extending the recommendations to funders and DOI registration agencies

The final section offers useful recommendations to publishers and funders. Publishers are encouraged to reassess and improve their practices for submitting funding metadata to Crossref. Funders are encouraged to monitor more systematically whether funded authors comply with grant conditions requiring funding to be acknowledged in publications.

I encourage the authors to extend this discussion. As they describe, funders often maintain detailed information about publications resulting from their funding. The discussion could therefore be broadened beyond improvements to the original publisher deposit workflow.

In particular, the authors could discuss the COMET (Collaborative Metadata) community approach to improving and enriching scholarly metadata, of which I am one of the organizers. The COMET model treats metadata completeness and quality as a shared responsibility. It enables trusted community partners to contribute metadata enrichments together with appropriate provenance information. This approach seems directly relevant to the present study because funders often possess rich internal data about the publications resulting from their funding. Funders such as NWO and FWF could make this data openly available as structured and properly documented metadata enrichments. These enrichments would complement the metadata provided by publishers. Just as publishers are expected to take responsibility for making comprehensive metadata available for the publications they publish, funders could contribute the information they maintain about the outputs resulting from their funding. This would complement, rather than replace, efforts to improve publisher workflows.

An essential element of the COMET model is round-tripping. Community-provided metadata enrichments should flow back into the infrastructures that maintain and disseminate DOI metadata, rather than remaining in disconnected downstream databases. This reduces fragmentation and allows a broad range of downstream systems and users to benefit from the enrichments. COMET and DataCite are already experimenting with such an approach. DataCite maintains enrichment records separately from the metadata provided by the original depositors. These records include provenance information, are subject to validation, and can be used to provide an enriched DOI record without altering the original record.

Another promising recommendation for Crossref and DataCite would be to enable publishers to deposit raw funding acknowledgement text, even when they are unable to transform it into structured funding metadata themselves, as also recommended by Mugabushaka et al. (2022). Publishers may be willing to share funding information but may lack the resources needed to identify funder names accurately and match them to funder IDs. In such cases, making the raw funding text openly available would already be a highly valuable contribution and would be far preferable to losing the information entirely. Crossref, DataCite, or trusted community partners could then take responsibility for extracting funder names, matching them to persistent identifiers, and identifying grant numbers. Retaining the original funding text alongside any derived metadata, with transparent provenance for both, would make it possible to verify and improve these enrichments over time. This would offer a pragmatic and scalable division of responsibilities: publishers make the information they possess available, while shared infrastructures and trusted community partners provide the specialist matching and enrichment capabilities.

I therefore encourage the authors to broaden their discussion and recommendations to address the complementary responsibilities of publishers, funders, DOI registration agencies such as Crossref and DataCite, and trusted community partners. This broader ecosystem perspective would strengthen the policy relevance of the paper. It recognizes that improving funding metadata should be a collective effort and should not depend exclusively on a single actor or a single point in the scholarly communication workflow.

Data and code

It is commendable that the authors have made available on Zenodo the data they are legally permitted to share, together with the source code used for DOI cleaning, Crossref API retrieval, and scraping of publisher landing pages. This is an important strength of the study and demonstrates a strong commitment to openness and transparency. The paper also clearly explains that the funding acknowledgement texts retrieved from Scopus and Dimensions cannot be shared because they come from proprietary sources.

The source code of the scripts has been made available in an RTF document. As a minor recommendation, I suggest making the scripts available as Google Apps Script code files using the .gs (Google Script) extension rather than as formatted text in an RTF file. The authors could additionally make the scripts available in a version-controlled code-sharing environment such as GitHub or GitLab. This would make the code easier to inspect, execute, improve, and reuse.

Competing interests: None.

Leave a comment