1. Introduction
Governments are under increasing pressure to ensure that public investment in research and innovation is efficient, accountable, and aligned with strategic objectives. This has led to the expansion of monitoring and evaluation practices aimed at assessing how funding is allocated and how it translates into scientific outputs. In this context, funding agencies are expected not only to allocate resources but also to maintain high-quality information systems that systematically capture data on funded projects, outputs, and impacts (Clements et al. 2017; Grassano et al. 2017; Mugabushaka et al. 2021). As research funding has become increasingly fragmented across multiple funders and sectors, the need for coordination has intensified, highlighting the role of centralised data warehouses in aggregating and harmonising funding information to support effective governance and analysis of public research investment. Such infrastructures have become increasingly important with the projectification of research, whereby funding is organised around discrete, time-bound projects rather than institutional block grants (Hodgson et al. 2019).
Systematic records of research funding provide several benefits, particularly for the transparent monitoring and evaluation of funding programmes and for enhancing accountability in public research investment. By enabling the measurement and tracking of how research investments are translated into scientific outputs (Clements et al. 2017; Rigby 2011; Thelwall et al. 2023), funding records inform allocation decisions and portfolio management, including the identification of inefficiencies such as redundant or overlapping investments (Ebadi & Schiffauerova 2016; Rigby & Julian 2014). More broadly, they enhance accountability by clarifying the links between financial inputs and research outputs, thereby demonstrating whether public investments address societal needs and policy objectives (Álvarez-Bornstein & Montesi 2021; MacLean et al. 1998). As funding information can also uncover relationships not observable through bibliographic indicators alone, such as co-authorship and citation patterns (El-Ouahi 2024), it has received increasing attention as a complementary instrument for analysing research outputs.
At the same time, the research information landscape has become increasingly distributed. Funding information is recorded across multiple heterogeneous data sources, including funders’ internal databases, bibliographic databases, and publisher-managed metadata systems, with the result that funding metadata are often fragmented, inconsistent, and only partially interoperable across systems (Álvarez-Bornstein et al. 2017; Mugabushaka 2021; Sirtes 2013). This fragmentation poses a fundamental challenge for funding agencies. The analytical value of funding information depends not only on its internal completeness and accuracy but also on its ability to interoperate with external data sources. Linking funders’ databases with bibliographic sources enables the enrichment of funding records with detailed bibliographic information and allows for the cross-validation of reported outputs.
However, such linkage is frequently constrained by incomplete identifier coverage, heterogeneous metadata standards, and inconsistencies in how funding information is recorded and curated across systems.
Against this backdrop, I investigate the interoperability, metadata quality, and coverage of funded publication outputs recorded in funders’ databases, using South Korea’s National Science & Technology Information Service (NTIS) as a case study. NTIS compiles and centralises information on public research funding from the South Korean government and is operated by the Korea Institute of Science & Technology Evaluation and Planning (KISTEP) under the Ministry of Science and ICT. NTIS annually receives reports from funding agencies on projects supported through all national research and development (R&D) programmes and their associated outputs, registering this information in a unified data schema to support decision-making and management of public R&D investment, including the compilation of national science and technology statistics, the allocation and adjustment of R&D budgets, and the planning and evaluation of national R&D programmes (Ministry of Science and ICT 2019). As a centralised funders’ database that integrates information across agencies in a standardised format, NTIS offers a useful case for examining challenges that are likely to confront funding information systems seeking to interoperate with bibliographic sources.
Unlike databases that operate at the level of individual funders, NTIS addresses data fragmentation by enabling the central government to collect and integrate funding information across agencies in a standardised format, thereby supporting a comprehensive overview of national research funding activities (Korean Government Ministries Concerned 2005). In addition, NTIS records are also publicly accessible, allowing users to view information on government R&D investment. These features make NTIS a well-suited case for examining how funders’ databases relate to bibliographic sources.
Three research questions are examined. First, to what extent are NTIS records linked to corresponding documents in bibliographic databases? Second, what is the quality of metadata related to publication outputs and their associated funding information in NTIS records? Third, how consistent is the coverage of Korean-funded publications between NTIS and other bibliographic databases? These questions are addressed empirically using NTIS, but the findings are intended to inform the development and governance of funding information systems more broadly. In particular, the study provides evidence on the strengths and limitations of funder-provided data relative to bibliographic sources, highlights how differences in data collection mechanisms shape the coverage of funding metadata, and demonstrates the extent to which combining multiple data sources can provide a more comprehensive representation of funded research outputs.
The remainder of the paper is structured as follows. The next section outlines the conceptual background and key issues addressed in this study, focusing on the interoperability of research information systems, metadata quality, and the coverage of Korean-funded publications across data sources. Sections 3, 4 and 5 address the first issue: Section 3 describes the data sources used in this study, and Section 4 explains the matching procedure employed to link NTIS records with bibliographic databases. Section 5 evaluates the reliability of the matching results. To address the second issue, Section 6 assesses the quality of metadata in NTIS records and identifies its limitations. Regarding the third issue, Section 7 compares funding information across different data sources, highlighting commonalities, differences, and potential complementarities. The final section concludes the paper with policy suggestions aimed at improving NTIS and, more broadly, at enhancing the interoperability and completeness of funders’ databases.
2. Three issues related to NTIS funding information
2.1 Interoperability of research information systems
The first research question concerns the interoperability of research information systems. The availability and transparency of funding information cannot be achieved simply by making data publicly accessible. As the scale of the research system and investment within the system expands, the research landscape is also becoming more complex and granular, with different research entities curating information in diverse contexts. Integrating and enriching research information through the linkage of associated objects enables fine-grained mapping of scientific activities and outputs, which in turn requires a high level of interoperability among data sources.
A challenge is how to interconnect all information that belongs together. Globally unique and persistent identifiers, such as DOIs, facilitate unambiguous and sustainable identification and referencing of relevant information. However, distributed local databases often lack sufficient DOI coverage (Babini et al. 2024). Furthermore, databases specialising in distinct types of entities require additional specific metadata, functions, and structural models, and adopt dedicated identifier systems to satisfy these requirements (Dappert et al. 2017). The heterogeneity of the identifier landscape, stemming from the differing natures and purposes of databases, constrains interconnection across them. Consequently, building a comprehensive research information system requires substantial effort to collect information on records missing persistent identifiers from multiple sources, such as indexing services and search engines. As a result, even when all relevant information exists, it is not uncommon for it to remain unconnected, leaving relationships between research objects unrepresented.
As discussed later, only 76.8% of NTIS records include DOIs, and no other identifiers linkable to bibliographic databases are provided. Consequently, while it is possible to identify the documents represented by these records through manual searches of unstructured output descriptions in bibliographic databases or search engines, interoperability across the entire dataset is not feasible. Although several studies have used NTIS records to measure the performance and productivity of Korea’s national R&D programmes (e.g. Jang 2025; D. Kim & Hwang 2022; M. Kim et al. 2025; Kwon & Kwon 2019; J. Lee et al. 2024; K. Lee et al. 2021), they have not provided sufficient methodological detail on how NTIS records lacking persistent identifiers were disambiguated into unique documents or how corresponding bibliographic information was acquired. To address this gap, the present study provides an explicit procedure for matching NTIS records with corresponding documents in bibliographic databases and examines the extent to which such linkage is feasible.
2.2 Metadata quality
Because scientometric analyses rely heavily on information such as authorship and citation data retrieved from bibliographic databases, the ability to obtain adequate, correct, and relevant information depends critically on the quality of the underlying metadata (Céspedes et al. 2025; Choudhury et al. 2023). In particular, linking records that lack persistent identifiers to bibliographic databases requires accurate descriptive metadata—such as titles and source information—that can reliably identify documents. Beyond enabling discoverability, interoperability, and accessibility, metadata quality also underpins the trustworthiness of a repository, shaping its perceived usefulness and credibility (Fear & Donaldson 2012; Frank 2022; Lin et al. 2020; Reiche & Hofig 2013). This raises the question of whether the metadata recorded in NTIS is of sufficient quality to support reliable linkage to bibliographic databases and, by extension, to ensure the credibility of analyses based on NTIS data.
I therefore assess the quality of selected metadata in NTIS records, focusing on bibliographic information—including DOI, title, and source—and the contribution rate attribute, which indicates the extent to which funding contributes to reported outputs and is used in compiling national science and technology statistics. While multiple criteria exist for evaluating (meta)data quality, two commonly used dimensions are adopted: completeness, defined as the extent to which data are present and of sufficient breadth, depth, and scope for the analytical task; and accuracy, defined as the extent to which data are correct and reliable (Cichy & Rass 2019; Pipino et al. 2002). Through this assessment, the study identifies key challenges that must be addressed to improve funding analysis based on NTIS records.
2.3 Coverage of Korean-funded publications by different data sources
One of the major challenges in analysing funding information is that funding metadata across data sources exhibits several limitations, including unstructured and heterogeneous descriptions (Álvarez-Bornstein et al. 2017; Grassano et al. 2017; Morillo & Álvarez-Bornstein 2018), incomplete reporting of funding information (Butler 2001; Costas & Yegros-Yegros 2013; Grassano et al. 2017; Koier & Horlings 2015), variations in acknowledgement practices influenced by field, journal, or regional context (Costas & van Leeuwen 2012; El-Ouahi 2024; Paul-Hus et al. 2017; Rigby 2011), and inconsistencies, errors, or loss of information during data processing (Álvarez-Bornstein et al. 2017; Koier & Horlings 2015; Liu 2020). Even when the same funding statement is present in publications, bibliographic databases may extract and structure the corresponding metadata differently (Gibson et al. 2022).
These limitations have motivated a growing body of research comparing funding information across data sources. On the one hand, studies have conducted comparisons among bibliographic databases themselves (Grassano et al. 2017; Kokol 2023; Kokol & Blažun Vošner 2018; Mugabushaka et al. 2022; Schares 2024). On the other hand, funding metadata from bibliographic databases has been compared with information maintained by funding bodies (Campbell et al. 2010; Ihli 2017; Kramer & de Jonge 2022; Mugabushaka et al. 2021; Powell 2019). Such comparisons allow researchers to assess the extent to which funding information from different databases—often characterised by varying degrees of transparency in their sources and collection methods—overlaps or complements one another (Mugabushaka 2021). The discrepancies identified across data sources suggest that reliance on a single database may introduce systematic bias into funding analyses; accordingly, multiple data sources should be jointly considered to obtain more robust and reliable results (Kokol 2023; Kokol & Blažun Vošner 2018; Powell 2019).
In this context, analysing public research funding in Korea similarly requires identifying both commonalities and differences across data sources. Although NTIS has established a comprehensive database of funding allocated through Korea’s national R&D programmes, the funding information compiled by funders may differ from that inferred from published research outputs. Moreover, because NTIS is limited to Korean public research funding, it cannot capture funding configurations involving other non-governmental sources that may support the same publications. This limitation constrains understanding of the broader funding landscape shaped by heterogeneous funding contributions from both within and outside the Korean government. Therefore, comparing funding information across data sources not only enables an assessment of the quality and coverage of each data source but also facilitates a more comprehensive understanding of funding dynamics through the complementarity of multiple data sources.
3. Data sources
3.1 National Science & Technology Information Service (NTIS)
To investigate publications that received public research funding in Korea and are registered with the National Science & Technology Information Service (NTIS), 1,111,175 journal article records reported as outputs of national R&D projects were retrieved from the NTIS database. NTIS provides this information in Excel format for articles published between 2007 and 2023 in journals indexed in the Science Citation Index Expanded (SCIE), including the following bibliographic details: a unique record identifier assigned to each funder report, project identifier, DOI, document title, journal title, and volume, issue, and page numbers. However, despite representing SCIE-indexed publications, the records lack identifiers that enable direct linkage to bibliographic databases, such as Web of Science accession numbers (UT). Furthermore, while DOI coverage has improved owing to recent reporting requirements, 23.2% of the records still lack DOIs.
Moreover, NTIS records represent concatenated project–output linkages rather than unique publications; a single article funded by multiple projects generates multiple NTIS records, as NTIS creates a separate entry for each funder report. DOIs provide the only reliable means of identifying identical publications across records; however, their limited coverage and the absence of explicit indicators clarifying whether different reports refer to the same document constrain accurate matching. This lack of interoperability restricts linkages with bibliographic databases and access to enriched metadata (e.g., ORCID, ROR). I address these limitations by linking NTIS local identifiers to global bibliographic identifiers.
3.2 Web of Science
Web of Science (WoS) was the primary target for metadata comparison with NTIS records, as NTIS covers only articles published in SCIE-indexed journals. The Week 13, 2025 snapshot of the CWTS in-house version of WoS was used, which includes the SCIE, Social Sciences Citation Index (SSCI), Arts & Humanities Citation Index (AHCI), and Conference Proceedings Citation Index (CPCI). Although NTIS records are restricted to SCIE-indexed publications, some records may also correspond to publications indexed in other WoS collections as a result of data curation processes.
Funding metadata from WoS were used to identify Korean-funded publications and to compare them with NTIS records. The WoS Core Collection extracts funding metadata from the acknowledgement sections of publications, including funding text, funding agency names, and grant numbers. However, the original WoS database does not provide detailed organisational attributes such as funder nationality, which limits its capacity to comprehensively identify Korean-funded research outputs. To address this limitation, I relied on the CWTS in-house funding database, in which funder organisational information has been standardised and enriched as follows. CWTS constructed a thesaurus of funding body and scheme names, and by mapping funding metadata indexed in WoS to thesaurus entries, inconsistently described funder names were linked to standardised organisational information. Furthermore, a dataset was constructed that includes hierarchical relationships among funders, nationality, and organisational type, offering a richer characterisation of funding sources identified and cleaned from funding text (van Honk et al. 2016).
3.3 OpenAlex
Because NTIS provides records only for publications in SCIE-indexed journals, linking to global identifiers and obtaining detailed metadata is primarily feasible through matching with WoS. Nevertheless, I also matched NTIS records with OpenAlex (CWTS in-house version, August 2025) due to its openness, interoperability, and extensibility with diverse databases such as Crossref, ORCID, and ROR, as well as its broader coverage—particularly of non-English and local publications (Céspedes et al. 2025; Chavarro et al. 2018; Culbert et al. 2025; International Science Council 2021; Scheidsteger et al. 2018; Scheidsteger & Haunschild 2022). As OpenAlex includes extensive coverage of Korean local publications, linking NTIS with OpenAlex enables scalability for analysing nationally relevant domains, such as citations from local publications and collaborations with local actors (Eum et al. 2025).
As with WoS, I used OpenAlex funding metadata—which includes funding text, funding agency names, and grant numbers—to compare and complement NTIS records and information on Korean-funded publications. Unlike WoS, which primarily extracts funding metadata from publication text, OpenAlex funding metadata is mainly derived from Crossref, which has collected funding information through publisher deposits since the launch of the Open Funder Registry (formerly FundRef) in 2013 (Meddings 2013). OpenAlex funding metadata is freely available, interoperable with different databases, and curated through an open and transparent process, offering a clear advantage for tracking funded outputs and their subsequent use. Although OpenAlex has recently been integrating grant metadata extracted from full texts or provided by funders into its database (Demes 2026), such metadata is not included in the version of the database used in this study.
4. Matching NTIS records with bibliographic databases
4.1 Matching attributes
NTIS records were matched to corresponding documents in WoS and OpenAlex by comparing the following bibliographic attributes: DOI, document title, source (ISSN and journal title), publication year, volume, issue, beginning page number, and article number. Although author names are included in NTIS records, their data quality was insufficient for reliable metadata comparison; therefore, they were excluded from the matching process and used only for manual verification of the sample. In addition, because the CWTS in-house version of OpenAlex does not provide article number information, this attribute was excluded from comparisons between NTIS and OpenAlex. Prior to matching, preprocessing was applied to all fields to remove diacritics, non-Roman characters, unnecessary prefixes, and non-numeric characters from numeric fields.
4.2 Matching procedure
The matching procedure was implemented through four consecutive steps (Figure 1).
At each step, a weighted matching score was calculated for candidate pairs, and only pairs exceeding a predefined threshold were retained. Technical details, including the attributes used to calculate the matching score and their corresponding weights, are provided in Appendix A. Once matched at a given step, NTIS records were excluded from subsequent steps, and duplicate matches were resolved by retaining the highest-scoring pair.

Figure 1. Flowchart of the matching procedure
The first step involved comparing NTIS records with WoS and OpenAlex records that contained DOIs. Because DOIs are unique and persistent identifiers across bibliographic databases, they served as the primary matching criterion. However, as incorrectly reported DOIs may link NTIS records to unrelated documents, consistency across additional matching attributes was also verified to ensure the validity of the matches.
The second step compared records sharing the same ISSN. Since ISSNs uniquely identify publication sources, they were used to restrict the set of candidate documents. For each NTIS record, candidate pairs were formed with all documents having matching ISSNs, and matching scores were calculated. The candidate pair with the highest score exceeding the threshold was selected. A discrepancy of up to two years between NTIS records and candidate documents was permitted.
The third step extended the approach used in the second step to journal name variants, such as abbreviated forms that omit articles, conjunctions, and prepositions, by matching them against the inventory of journal name variants provided by WoS and OpenAlex. As in the second step, a discrepancy of up to two years was permitted, and the candidate pair with the highest score exceeding the threshold was selected.
In the final step, all remaining unmatched NTIS records were compared against all available bibliographic documents using the full set of matching attributes. Records that could not be matched through any of the four steps were classified as unmatched.
This multi-step procedure was designed to maximise successful matches while reducing computational load by progressively narrowing the candidate sets. Except for DOI-based matching in the first step, the remaining steps require score calculations for all candidate pairs, which can be computationally intensive. Nonetheless, because NTIS records are registered only after manual verification of reported publications and are limited to articles published in SCIE-indexed journals (Korea Institute of Science and Technology Evaluation and Planning 2023), all records should in principle correspond to documents indexed in WoS and OpenAlex, assuming accurate metadata.
Applying this matching procedure yielded 1,109,458 and 1,103,829 matched document pairs from WoS and OpenAlex, corresponding to 581,909 and 578,166 distinct documents, respectively. Figure 2 illustrates the number of NTIS records remaining after each step. In the first step, 75.6% of records were matched in WoS and 75.4% in OpenAlex, substantially reducing the computational burden for subsequent steps.

Figure 2. Number of remaining NTIS records after each step in the matching procedure
5. Verification of matching results
5.1 Quantitative assessment of matched and unmatched records
To assess whether NTIS records were correctly linked to WoS and OpenAlex, matches were cross-validated by examining whether multiple independent metadata attributes from the corresponding WoS and OpenAlex documents referred to the same publication.
When the linked documents indicated the same publication, the match was considered correct. Following the approach of Visser et al. (2021), matching scores were calculated to assess metadata consistency for NTIS records matched in both databases. Of the 1,111,175 NTIS records, 1,102,944 matched to both databases, of which 1,100,908 were confirmed as correct matches. Cross-validation was not conducted for records matched to only one database. The overall matching outcomes are summarised in Figure 3.

Figure 3. The range of NTIS records matched with WoS and OpenAlex
Precision and recall were then calculated to assess the reliability of the matching results. Precision reflects the proportion of correctly matched results and is defined as the number of correct matches (true positives) divided by the sum of the number of correct matches and the number of incorrect matches (false positives), i.e. the number of all matches. Because no ground-truth dataset was available, cross-validated matches between WoS and OpenAlex were used as proxies for correct matches. Based on this approach, precision was estimated at 99.23% for WoS and 99.74% for OpenAlex.
Recall measures the extent to which corresponding documents across databases are successfully matched and is defined as the number of correct matches divided by the sum of correct matches and missed matches (false negatives). To estimate recall, I manually reviewed the bibliographic information of random samples of unmatched NTIS records: 92 records from the 1,717 NTIS records unmatched to WoS and 95 records from the 7,346 NTIS records unmatched to OpenAlex (95% confidence level with a 10% margin of error). Among these samples, 7 and 9 records, respectively, were identified as missed matches. The total number of missed matches was then estimated by extrapolating these proportions to all unmatched records. Using this method, 131 and 696 records were estimated as missed matches for WoS and OpenAlex, respectively, resulting in recall estimates of 99.99% for WoS and 99.94% for OpenAlex.
While there is often a trade-off between precision and recall, both metrics exhibited very high values in this study. This outcome reflects the nature of the dataset: NTIS records are registered only after verification by funding agencies and NTIS data administrators, ensuring that corresponding publications exist and that the associated metadata are sufficiently accurate to support reliable matching.
5.2 Qualitative assessment of incorrect and missed matches
Because the calculation of precision and recall relies on estimated values for true positives and false negatives, the limitations of the estimation approach warrant careful consideration. Several issues arise when estimating true positives through metadata comparison between WoS and OpenAlex. First, NTIS records matched to only one database were excluded from cross-validation, even when those matches were correct. Given differences in database coverage (Visser et al. 2021), not all funded publications may have been covered by both WoS and OpenAlex. Second, technical judgments of correctness based on metadata consistency may not fully reflect actual document correspondence. If WoS or OpenAlex contains erroneous metadata, empirically correct matches may be misclassified as incorrect. Conversely, NTIS records linked to incorrect documents in both WoS and OpenAlex may be classified as correct matches if the incorrect records share consistent metadata across the two databases.
To address these concerns, I conducted a manual review of bibliographic information for random samples of NTIS records: 97 records were drawn from the 1,100,908 NTIS records classified as correct matches to evaluate true positives. To assess false positives, 95 records were selected from the 7,399 NTIS records matched to only one database, and 92 records were sampled from the 2,036 NTIS records classified as incorrect matches through cross-validation between WoS and OpenAlex (95% confidence level with a 10% margin of error). No errors were identified among the sampled correct matches. Similarly, no metadata inconsistencies or incorrect linkages were found among NTIS records matched to only WoS or only OpenAlex. In addition, 56 records were identified as having no corresponding documents in either WoS or OpenAlex; many of these publications appeared in journals that were previously indexed by WoS but were subsequently delisted.
Among the records classified as incorrect matches through cross-validation, 80 exhibited metadata inconsistencies. In many cases, document titles varied in their representation of symbols related to constants, variables, and physical quantities, including Greek characters. Such metadata quality issues—particularly prevalent in OpenAlex—have been documented in prior studies (Alperin et al. 2024; Borrego & Urbano 2025; Céspedes et al. 2025; Culbert et al. 2025; Delgado-Quirós & Ortega 2024; Haupka et al. 2025; Mongeon et al. 2023; Zhang et al. 2024). In light of these findings, mismatches identified by the methodology cannot be attributed solely to deficiencies in the matching procedure; they may also reflect incomplete or erroneous metadata in the underlying data sources.
Given that NTIS records undergo human verification, the presence of unmatched records necessitates further examination of the matching process. To justify the estimation of false negatives, I therefore investigated the aforementioned random samples of unmatched records. Several records classified as true negatives—those for which corresponding documents could not be identified—exhibited the following issues. First, 98 records contained substantial typographical errors in publication titles or journal names, in some cases involving Korean characters. Second, four records corresponded to book chapters, which fall outside the document coverage of WoS and OpenAlex. Third, consistent with findings from the incorrect matches, 60 records were associated with missing or incomplete metadata in WoS or OpenAlex, particularly inaccurate document titles. Although manual verification confirmed the existence of these publications, algorithmic matching at scale could not identify them. Accordingly, these records were not attributed to deficiencies in the matching procedure and were classified as true negatives.
However, 7 and 9 records for WoS and OpenAlex failed to match corresponding documents despite the absence of the issues described above. In these cases, algorithmic matching was theoretically feasible, but high Levenshtein distances between document titles prevented matching scores from exceeding the threshold. Sometimes minor variations in character representation led to increased string distances. These records could potentially be matched through additional string cleaning or by assigning lower weights to title differences. Consequently, these cases were classified as missed matches attributable to the matching procedure.
6. Assessment of the quality and limitations of NTIS metadata
6.1 Quality assessment of bibliographic metadata in NTIS
As with bibliographic databases, metadata in NTIS records may be incomplete despite undergoing human verification. To address the second research question, this section compares the metadata of NTIS records with those of matched WoS and OpenAlex documents in order to assess the completeness and accuracy of metadata for funded publications recorded in NTIS. The comparison was restricted to the attributes used in the matching process with WoS and OpenAlex: DOI, document title, ISSN, journal name, publication year, volume, issue, beginning page number, and article number.
Completeness refers to the extent to which metadata fields are filled in, whereas accuracy refers to the degree of agreement between corresponding metadata values across document pairs. Because metadata may be missing not only in NTIS but also in WoS and OpenAlex, accuracy was calculated only for document pairs in which the attribute under comparison was present in both records. Title accuracy was assessed using the mean and standard error of the Levenshtein distance between corresponding titles. Since metadata for the same document may differ across bibliographic databases (Crotty 2024), the consistency of metadata between WoS and OpenAlex was also assessed. As summarised in Table 1, metadata for all attributes except article number exhibited high levels of completeness and accuracy, with strong consistency observed between corresponding documents in WoS and OpenAlex.
Table 1. Completeness and accuracy of bibliographic metadata in NTIS
| Attributes |
NTIS |
NTIS-WoS |
NTIS-OpenAlex |
WoS-OpenAlex |
| Assessment |
Completeness |
Accuracy |
Accuracy |
Consistency |
| Records |
1,111,175 |
1,109,458 |
1,103,829 |
577,489 |
| DOI |
76.8% |
98.6% |
98.7% |
99.7% |
| Title |
100% |
0.0147 |
0.0250 |
0.0141 |
| ISSN |
92.4% |
99.4% |
99.5% |
99.1% |
| Journal name |
100% |
98.1% |
86.4% |
87.1% |
| Year |
100% |
99.9% |
83.9% |
82.8% |
| Volume |
99.4% |
98.9% |
97.8% |
99.1% |
| Issue |
86.5% |
96.4% |
92.3% |
98.4% |
| Beginning page |
78.0% |
94.6% |
88.2% |
99.2% |
| Article number |
40.6% |
20.2% |
– |
– |
Note: Records in the WoS-OpenAlex entry indicate the number of unique pairs of corresponding documents.
6.2 Limitations of funder records in measuring the fractional credit of projects
NTIS records include a metadata element called the “contribution rate,” which is not provided by bibliographic databases. For each project–publication pair, a separate NTIS record is created in which the project’s share of credit for the publication is recorded. Figure 4 illustrates the assignment of contribution rates. Publication X is reported exclusively as the output of Project A and therefore receives full (100%) credit. Publication Y is reported as a joint output of Projects A and B from different funders, with NTIS data administrators assigning contribution rates of 40% and 60%, respectively. In this way, the contribution rate functions as a weight for the fractional counting of funded publications. Because it is assigned centrally by the government agency responsible for all publicly funded projects, partial credit can be identified at the project level. Furthermore, by aggregating contribution rates across records, the total number of funded publications can be counted, a figure that is also used in national science and technology statistics (see Appendix B). This feature helps address a key limitation of acknowledgement-based funding metadata, which often lacks detail on the relative scale of contributions when multiple funders are involved (Rigby 2011). In this respect, contribution rate metadata serves as a “value-added language” (Greenberg 2017).

Figure 4. A simplified example of reporting on publication output and contribution rates for projects
However, I identified several potential sources of error in statistics derived from aggregating contribution rates, as summarised in Table 2. First, NTIS intentionally omits contribution rate values for 21.0% of records, as illustrated by the link between Project D and Publication Z in Figure 4. According to KISTEP, which operates NTIS as the validation authority, when the same output from a single project is reported multiple times, credit is assigned to only one report and the remaining records are left blank (Korea Institute of Science and Technology Evaluation and Planning 2023). However, duplicate project–publication pairs were relatively rare (346 in WoS; 388 in OpenAlex) compared with the overall number of records with missing contribution rates, suggesting that KISTEP has also exercised discretionary judgment in determining that certain projects did not warrant assigned credit. Such adjustments may result in underestimation of performance at the programme or project level. Moreover, even when partial credit is assigned, it has frequently been done without consultation with funding agencies, leading to persistent objections from funders (Korea Institute of Science and Technology Evaluation and Planning 2017). The contribution rate should therefore be interpreted as an administratively assigned indicator of funding involvement rather than a precise, empirically grounded measure of actual contribution.
Table 2. Issues with NTIS contribution rate metadata
|
WoS |
OpenAlex |
| 1. Missing contribution rate |
| NTIS records without credit |
233,754 (21.0%) |
| Duplicate project-publication pairs |
346 |
388 |
| 2. Aggregated contribution rates exceeding 100% |
| Documents exceeding contribution rates total of 100% |
1,632 (0.28%) |
1,761 (0.30%) |
| Surplus of aggregated contribution rates |
1643.9 |
1809.5 |
| 3. Aggregated contribution rates of 100% with external funding sources |
| Documents with contribution rates total of 100% |
573,795 (98.6%) |
569,483 (98.5%) |
| Documents with non-NTIS funding metadata |
446,093 (76.7%) |
205,948 (35.6%) |
| Non-NTIS funding (non-Korean funding) |
1,331,601 (383,721) |
459,347 (140,851) |
Second, instances of overcounting persist even under fractional counting. By definition, the sum of contribution rates across all projects associated with a single publication should not exceed 100%, but in practice the sum of contribution rates does sometimes exceed 100%, as illustrated for Publication Z in Figure 4. 1,632 and 1,761 documents matched to WoS and OpenAlex, respectively, exhibited aggregated contribution rates exceeding this threshold. Although these cases represent a small fraction of total funded publications, they indicate that manual assignment of contribution rates does not fully prevent double counting.
Third, even when aggregated contribution rates do not exceed 100%, further limitations remain. NTIS adjusts credit only for projects funded by Korean public agencies, without accounting for contributions from industry, foreign funders, or other non-governmental sources. Consequently, for publications assigned a total contribution rate of 100% in NTIS that also received external funding, the actual share of Korean public research funding is necessarily lower than reported. To examine this, I identified additional funding information for 573,795 and 569,483 matched publications in WoS and OpenAlex, respectively, that were assigned a total NTIS contribution rate of 100%. An additional 1,331,601 and 459,349 funding metadata elements were retrieved from 446,093 and 205,948 documents in WoS and OpenAlex, respectively, of which 383,721 and 140,851 represented non-Korean funding sources. For such cases, the actual contribution of Korean public research funding is below 100%, a discrepancy that cannot be identified from NTIS alone.
In summary, while contribution rate metadata is a valuable feature of NTIS not available in bibliographic databases, its reliability is constrained by missing values, occasional overcounting, and the exclusion of non-Korean funding sources, all of which may lead to an overstatement of Korean public research funding in national statistics.
7. Comparison of Korean-funded publications across different data sources
7.1 Comparison of coverage of Korean-funded publications
Consistent with prior literature documenting variation in the coverage of funded publications across data sources, funding metadata from WoS and OpenAlex identify Korean-funded publications that are not registered in NTIS, despite NTIS’s role as the central repository for Korean public research funding. This observation motivates the third research question: to what extent does each data source cover Korean-funded publications? This section addresses this question by comparing Korean-funded publications and associated funding information retrieved from WoS and OpenAlex with those recorded in NTIS.
To retrieve Korean-funded publications, the following criteria were applied. First, Korean funders had to be identified in the funding metadata of publications in WoS and OpenAlex. Second, to ensure comparability with NTIS records, the journal scope was restricted to journals in which NTIS-reported publications had appeared. Third, because NTIS provides information only for publications reported between 2007 and 2023, the publication year was limited to this period. Fourth, funding metadata entries with blank funding identifier—where only funder names were acknowledged, and no detailed information (e.g., project titles) was provided—were also included. Fifth, to determine whether publications received public research funding, both strict and loose criteria were applied to classify funder types.
Under the strict criteria, only government and funding bodies were selected (i.e., “funding organisation” and “governmental institution” types in WoS, and “government” and “funder” types in OpenAlex). The loose criteria expanded the scope to account for cases in which organisations were incorrectly typed or functioned as funding intermediaries despite not being classified as government or funding bodies. This distinction is particularly relevant in Korea, where government-funded research institutes also function as funding agencies.
Accordingly, “research organisation” and “university” types were additionally included for WoS, while “education” and “facility” types were included for OpenAlex. Because OpenAlex does not directly provide organisational type information for funders, Research Organization Registry (ROR) identifiers were used to link funders to organisational attributes in the ROR database.
Using funding information from NTIS, WoS, and OpenAlex, Korean-funded publications were retrieved as summarised in Table 3 and visualised in Figure 5. The results indicate that 77.4% of funded publications reported in NTIS were also retrieved from WoS or OpenAlex under the strict criteria, increasing to 79.5% under the loose criteria. Specifically, 73.0% of NTIS-funded publications were covered by WoS under the strict criteria (76.3% under the loose criteria). Approximately one-third of all retrieved publications were captured by only a single data source. The number of publications retrieved from OpenAlex was smaller than that retrieved from NTIS and WoS. However, while the difference between strict and loose criteria amounted to approximately 61,000 publications in WoS, almost no difference was observed in OpenAlex. Additional checks confirmed that a small number of NTIS publications were excluded due to errors in reported publication years, which led to their omission from coverage in WoS and OpenAlex; however, these cases did not affect the overall results.
Table 3. Coverage of Korean-funded publications by criteria for funder type
| Criteria |
NTIS |
WoS |
OpenAlex |
N∩W |
N∩O |
W∩O |
N∩W∩O |
N∪W∪O |
| Strict |
581,714
(81.9%) |
515,191
(72.5%) |
275,110
(38.7%) |
424,959
(59.8%) |
212,333
(29.9%) |
211,595
(29.8%) |
187,030
(26.3%) |
710,353 |
| Loose |
581,714
(78.4%) |
576,299
(77.7%) |
275,113
(37.1%) |
443,814
(59.8%) |
212,333
(28.6%) |
229,070
(30.9%) |
193,457
(26.1%) |
741,561 |
| Difference |
|
61,108 |
3 |
18,855 |
0 |
17,475 |
6,427 |
31,208 |
Note: Percentages are calculated relative to the total number of publications covered by at least one of the three data sources (N∪W∪O).

Figure 5. Coverage of Korean-funded publications (strict criteria)
Note: Percentages are calculated relative to the total number of publications covered by at least one of the three data sources.
7.2 Investigation of coverage differences across data sources
Differences in coverage among NTIS, WoS, and OpenAlex are primarily attributable to the timing of funding information collection. WoS began extracting funding information from SCIE-indexed publications in 2008, whereas Crossref—the primary source of funding metadata for OpenAlex—began collecting such information in 2013. Consequently, as shown in Figure 6, many publications reported in NTIS were not identified as Korean-funded publications in WoS or OpenAlex during earlier years. Although there was initially a substantial gap in coverage between WoS and OpenAlex, coverage has increased over time and the gap has gradually narrowed.

Figure 6. Temporal trend in the coverage of Korean-funded publications
Another reason of coverage differences lies in how funding information is collected. WoS builds its funding database by extracting funding text from acknowledgement sections of publications, whereas OpenAlex relies primarily on funding metadata deposited by publishers via Crossref. Because publishers differ in their metadata collection and deposition practices, the availability and quality of funding metadata vary accordingly: only some publishers provide submission systems that allow authors to select funders from a standardised list at the time of manuscript submission (Kramer & de Jonge 2022). Even when funding statements are prepared in a standardised format, publishers may omit details (de Jonge 2025), selectively deposit only certain metadata to Crossref, or withhold metadata they do not wish to make openly available (van Eck & Waltman 2025).
Figure 7 visualises the extent to which OpenAlex provides any funding metadata—regardless of funding source—for NTIS publications, disaggregated by publisher. Elsevier and Springer Nature, the publishers associated with the largest numbers of NTIS records and reported publications, provided funding metadata for only 55.8% and 23.6% of their publications, respectively. Among the 30 publishers with the highest publication counts, four provided no funding metadata at all: American Scientific Publishers, Institute of Electronics, Information and Communication Engineers, Mary Ann Liebert, Inc., and Spandidos Publications. These findings indicate that, irrespective of the completeness of reporting in NTIS, coverage in OpenAlex may be substantially constrained by publisher-level metadata practices.

Figure 7. Coverage of NTIS publications with OpenAlex funding metadata by publisher
The failure of some NTIS publications to be captured as Korean-funded in WoS or OpenAlex is also influenced by funding acknowledgement practices in Korea. Although all Korean R&D programmes require recipients to report their research outputs, not all programmes mandate the acknowledgement of funding sources in publications. For example, the Brain Korea 21 (BK21) programme—which accounts for 29.3% of all NTIS publication records—recommends but does not mandate funding acknowledgements, leaving acknowledgement decisions to researchers’ discretion. As a result, of the 172,456 NTIS records for which Korean public funding information was not identified in WoS or OpenAlex, 74,659 were outputs from BK21, representing 22.9% of the 325,321 records submitted under the BK21 programme. In addition, prior studies have noted the limited discussion and standardisation of acknowledgement norms in Korea (H. Lee et al. 2020).
Consequently, discrepancies may arise between the mandatory reporting of outputs to funders and the voluntary inclusion of funding acknowledgements in published articles.
7.3 Comparison of funding source information across data sources
For all Korean-funded publications retrieved from WoS and OpenAlex, I first examined their reported funding sources. Because the objective was to compare funding source information as comprehensively as possible across the two databases, loose criteria for funder type were applied to maximise inclusion. As expected, the majority of funders identified in the funding metadata of WoS and OpenAlex were Korean government organisations, as shown in Table 4. In addition, a substantial number of records listed foreign funding bodies, such as the National Natural Science Foundation of China and the National Science Foundation of the United States, reflecting the prevalence of internationally co-funded research.
Table 4. Major funders of retrieved Korean-funded publications (full counting)
| WoS |
OpenAlex |
| Funding body |
Country |
Publication |
Funding
records |
Funding body |
Country |
Publication |
Funding
records |
| National Research Foundation of Korea |
Korea |
327,726 |
495,055 |
National Research Foundation of Korea |
Korea |
169,932 |
222,865 |
| Government of South Korea |
Korea |
152,590 |
188,428 |
Ministry of Science, ICT and Future Planning |
Korea |
24,348 |
30,397 |
| Ministry of Trade, Industry and Energy, South Korea |
Korea |
67,007 |
87,973 |
Ministry of Education |
Korea |
21,613 |
24,876 |
| Ministry of Health & Welfare, South Korea |
Korea |
37,503 |
47,357 |
Ministry of Education, Science and Technology |
Korea |
19,326 |
23,135 |
| National Research Council of Science & Technology |
Korea |
31,290 |
35,275 |
Ministry of Trade, Industry and Energy |
Korea |
16,792 |
18,510 |
| Rural Development Administration |
Korea |
19,450 |
24,228 |
Ministry of Science and ICT, South Korea |
Korea |
10,452 |
12,685 |
| Korea Research Foundation |
Korea |
15,157 |
17,656 |
Korea Institute of Energy Technology Evaluation and Planning |
Korea |
9,870 |
10,553 |
| Korea Science and Engineering Foundation |
Korea |
13,754 |
15,092 |
Ministry of Health and Welfare |
Korea |
6,894 |
8,014 |
| National Natural Science Foundation of China |
China |
12,311 |
24,869 |
Korea Health Industry Development Institute |
Korea |
6,358 |
6,839 |
| National Science Foundation |
US |
11,659 |
18,488 |
National Natural Science Foundation of China |
China |
6,066 |
11,517 |
| Others |
373,017 |
486,159 |
Others |
232,383 |
260,118 |
| Total |
1,061,464 |
1,440,580 |
Total |
524,034 |
629,509 |
Despite this broad coverage, the accuracy and granularity of funder identification remain limited. In WoS funding metadata, “Government of South Korea” ranked as the second most frequently identified funder. While this designation clearly indicates public funding, it does not specify the responsible funding agency. This limitation arises from both the mechanisms used by WoS to extract funding information and historical funding acknowledgement practices in Korea. In earlier periods, many national R&D programmes required researchers to acknowledge funding using generic terms such as “Korean Government” (e.g. Korean Research Foundation 2009; Ministry of Education, Science, and Technology 2009). Although more recent guidelines require acknowledgement of specific funding agencies, WoS relies on textual extraction from published acknowledgements and therefore cannot retroactively disaggregate such generic funder designations for previously published articles.
To address this limitation, I sought to infer specific funding agencies by matching NTIS project identifiers with grant identifiers extracted from publications. When WoS grant identifier and NTIS project identifier matched, detailed funding agency information could be retrieved from the corresponding NTIS record, even when WoS classified the funder generically as the Government of South Korea. The matching procedure applied the following criteria. First, the nationality of the funding metadata in WoS had to be Korean.
Second, the publication year had to be the same as or later than the project’s funding year. Third, NTIS project identifiers consist of two components: a centrally assigned NTIS project identifier and an internal management identifier assigned by the funding agency or the recipient’s affiliated organisation. A match was considered valid if the WoS grant identifier corresponded to either of these components. Fourth, identifiers shorter than five characters were excluded; all centrally assigned NTIS project identifiers are 10 characters long, and shorter internal identifiers were deemed unreliable for comparison. Using this approach, 29,408 NTIS project identifiers were matched to 74,158 publications. The major funders identified through this approach are summarised in Table 5. These results demonstrate that NTIS can not only identify funded publications not captured by bibliographic databases but also enrich them with more detailed funding information. Where WoS and OpenAlex identify funders only at the ministry level, this approach can further disambiguate them to the level of specific agencies using funder registries.
Table 5. Major funders identified from WoS “Government of South Korea” sources via NTIS
| Funding body |
Publications |
Funding records |
| National Research Foundation of Korea |
54,261 |
18,794 |
| Korea Science and Engineering Foundation |
9,388 |
1,706 |
| Korea Planning & Evaluation Institute of Industrial Technology |
4,786 |
1,519 |
| Korea Institute of Energy Technology Evaluation and Planning |
2,682 |
787 |
| Korea Health Industry Development Institute |
2,590 |
688 |
| Institute of Information & Communications Technology Planning &
Evaluation |
1,510 |
411 |
| National Research Council of Science & Technology |
1,239 |
101 |
| Korea Institute for Advancement of Technology |
1,203 |
367 |
| Daegu Gyeongbuk Institute of Science & Technology |
951 |
255 |
| Korea Technology and information Promotion Agency for SMEs |
861 |
560 |
| Others |
11,342 |
4,922 |
| Total |
90,813 |
30,110 |
I then examined funding sources for Korean-funded publications retrieved from WoS and OpenAlex that were not registered in NTIS. Applying the same matching procedure, funding identifiers from WoS and OpenAlex were compared with NTIS project identifiers to identify potentially unreported NTIS-funded publications. Strict criteria were applied to exclude publications likely supported by intramural funding from universities or public research institutes. Using this method, 59,088 funding metadata from 25,004 WoS publications and 28,583 funding metadata from 10,956 OpenAlex publications were matched to NTIS project identifiers. Consequently, 27.7% of WoS publications and 15.7% of OpenAlex publications retrieved as Korean-funded but not registered in NTIS can be considered unreported NTIS-funded outputs.
The coverage of Korean-funded publications across NTIS, WoS, and OpenAlex can be summarised as follows. First, although there is substantial overlap among the data sources, coverage differs considerably, with approximately one-third of publications captured by only a single source. This finding indicates that reliance on a single data source may lead to incomplete identification of funded research outputs. Second, although OpenAlex coverage of funded publications has expanded over time, it depends entirely on publisher-deposited metadata, leaving many funded publications unrecorded due to publishers’ selective or incomplete metadata provision. Third, integrating funding information from multiple data sources enables more comprehensive funding analyses. NTIS provides detailed and authoritative information on Korean public research funding, including projects and agencies not fully captured in bibliographic databases. Meanwhile, WoS and OpenAlex offer additional coverage by identifying Korean-funded publications that were not reported to NTIS, as well as funding information from non-Korean or non-governmental sources.
Accordingly, a comprehensive analysis of Korean research funding configuration requires the combined use of multiple data sources rather than reliance on any single database.
8. Concluding remarks and policy suggestions
The following research questions were addressed regarding the database of funded publication outputs provided by NTIS: To what extent are NTIS records linked to corresponding documents in bibliographic databases? What is the quality of metadata related to publication outputs and their associated funding information in NTIS records? How consistent is the coverage of Korean-funded publications between NTIS and other bibliographic databases? Three tasks were carried out to address these questions: matching NTIS records with WoS and OpenAlex, assessing the metadata quality of NTIS records, and comparing the coverage of Korean-funded publications across the three data sources.
The main findings can be summarised as follows. First, more than 99% of NTIS records were linked to corresponding documents in WoS and OpenAlex. This high match rate is unsurprising, as NTIS provides information exclusively on funded publications in SCIE-indexed journals, whose existence is verified by funding agencies and data administrators.
Second, the bibliographic metadata in NTIS records demonstrates a high level of completeness and accuracy, largely due to manual verification. The main limitations arise from the contribution rates assigned to funding sources: NTIS intentionally omits these from some records and adjusts credit only among projects registered within NTIS, without accounting for non-Korean or non-governmental funding, which may lead to an overstatement of Korean public research funding in national statistics. Third, although there is substantial overlap in the coverage of Korean-funded publications across NTIS, WoS, and OpenAlex, approximately one-third of publications are captured by only a single data source. This gap reflects differences in how funding information is collected: NTIS relies on output reporting by funding recipients, WoS extracts funding acknowledgements from publication texts, and OpenAlex depends on metadata deposited by publishers. Such mechanism-driven coverage gaps can be a general challenge for any funding body that relies on researcher-submitted reports. However, these differences also create opportunities for complementarity, as WoS and OpenAlex reveal funded outputs not reported to funding agencies, while NTIS disambiguates funders generically identified in bibliographic databases and captures funding sources not acknowledged in publications but reported directly to funders.
Despite the high match rate and generally strong bibliographic metadata quality, there remains scope for improvement in data interoperability and inclusiveness. First, a more proactive metadata policy is needed. While most recent NTIS records include DOIs, older records could retrospectively incorporate them based on matched documents. Furthermore, domain-specific persistent identifiers such as ORCID and ROR are not currently applied to authors and affiliations, resulting in inconsistencies and limiting interoperability with external data sources. Given that many SCIE-indexed journals already support ORCID identifiers, requiring funded researchers to provide ORCID identifiers and standardised affiliation identifiers based on ROR would substantially enhance data usability. Second, the scope of publicly available funded publication information should be expanded beyond SCIE-indexed journals. Although funding agencies report all funded publications to NTIS, information on non-SCIE publications is not publicly accessible. Restricting public records to SCIE-indexed journals covers only a subset of science and engineering publications and excludes social science and humanities research as well as outputs in local journals indexed in the Korea Citation Index, leaving a substantial portion of national research output invisible. Expanding the publicly available dataset alongside improvements in metadata quality would enhance the comprehensiveness of the national research outputs.
I examined publication outputs of Korean public research funding at scale, but leaves room for more fine-grained analyses of funding metadata. While the coverage of Korean-funded publications was compared and unreported outputs and ambiguously identified funders were identified using NTIS project identifiers, differences in the linkage between publications and funding sources across data sources were not examined. Because NTIS, WoS, and OpenAlex each collect funding metadata through distinct mechanisms, not only the coverage of funded publications but also the identified funding sources for the same publications may differ. Future research could provide a more comprehensive understanding of funding–publication relationships by comparing coverage and consistency at the level of publication–funding pairs rather than publications alone.
From a policy perspective, these findings carry broader implications for funding agencies that manage research information. The effectiveness of funding information systems depends on the quality and interoperability of the data they provide, and limitations in metadata standardisation, identifier adoption, and data coverage constrain the integration and interpretation of funding information. As the value of funding databases lies not only in their coverage but in their capacity to serve as interoperable nodes within a wider research information ecosystem, two priorities are particularly salient. First, the systematic use of persistent identifiers, including retrospectively where feasible, substantially enhances the linkability of funding records with external data sources and reduces inconsistencies arising from unstructured manual entries. Second, expanding the scope of publicly available funding records beyond selectively indexed journals provides a more representative picture of national research output. More fundamentally, funding agencies should invest in the continuous governance, standardisation, and integration of their information systems to support persistent identifier adoption, data interoperability, and structured metadata submission. Such efforts will enable funding databases to function as interoperable components within an increasingly complex global research information ecosystem.
Data availability
Although the data used in this study can be obtained by submitting a request to NTIS, redistribution of all or part of the data to third parties is restricted. Therefore, NTIS funding records and the identifiers of the corresponding WoS and OpenAlex documents cannot be made publicly available. To ensure reproducibility, the code used to match NTIS records with WoS and OpenAlex is provided at the following link: https://github.com/eumsoohong/ntis_matching.
Acknowledgements
The author would like to thank Ludo Waltman and Ismael Rafols, whose comments and suggestions have contributed to improving the manuscript. The author is also grateful to Jinseo Park for his support in the smooth acquisition of NTIS records.
References
Alperin, J. P., Portenoy, J., Demes, K., Larivière, V., & Haustein, S. (2024). ‘An analysis of the suitability of OpenAlex for bibliometric analyses’,.
Álvarez-Bornstein, B., & Montesi, M. (2021). ‘Funding acknowledgements in scientific publications: A literature review’, Research Evaluation, 29/4: 469–88. DOI: 10.1093/reseval/rvaa038
Álvarez-Bornstein, B., Morillo, F., & Bordons, M. (2017). ‘Funding acknowledgments in the Web of Science: completeness and accuracy of collected data’, Scientometrics, 112/3: 1793–812. DOI: 10.1007/s11192-017-2453-4
Babini, D., Garcia, A. B., Costas, R., Matas, L., Rafols, I., & Rovelli, L. (2024). ‘Not only Op en, but also Diverse and Inclusive: Towards Decentralised and Federated Research Infor mation Sources’. DOI: 10.59350/gmrzb-e2p83
Borrego, Á., & Urbano, C. (2025). ‘OpenAlex: Features, advantages and limitations of an open database for retrieving and analysing scholarly outputs’,.
Butler, L. (2001). ‘Revisiting bibliometric issues using new empirical data’, Research Evaluation, 10/1: 59–65. DOI: 10.3152/147154401781777141
Campbell, D., Picard-Aitken, M., Côté, G., Caruso, J., Valentim, R., Edmonds, S., Williams, T., et al. (2010). ‘Bibliometrics as a Performance Measurement Tool for Research Ev aluation: The Case of Research Funded by the National Cancer Institute of Canada’, Am erican Journal of Evaluation, 31/1: 66–83. DOI: 10.1177/1098214009354774
Céspedes, L., Kozlowski, D., Pradier, C., Sainte-Marie, M. H., Shokida, N. S., Benz, P., Poitr as, C., et al. (2025). ‘Evaluating the linguistic coverage of OpenAlex: An assessment of metadata accuracy and completeness’, Journal of the Association for Information Science and Technology, 76/6: 884–95. DOI: 10.1002/asi.24979
Chavarro, D., Ràfols, I., & Tang, P. (2018). ‘To what extent is inclusion in the Web of Science an indicator of journal “quality”?’, Research Evaluation, 27/2: 106–18. DOI: 10.1093/reseval/rvy001
Choudhury, M. H., Salsabil, L., Jayanetti, H. R., Wu, J., Ingram, W. A., & Fox, E. A. (2023). ‘MetaEnhance: Metadata Quality Improvement for Electronic Theses and Dissertations of University Libraries’,.
Cichy, C., & Rass, S. (2019). ‘An Overview of Data Quality Frameworks’, IEEE Access, 7: 2 4634–48. DOI: 10.1109/ACCESS.2019.2899751
Clements, A., Reddick, G., Viney, I., McCutcheon, V., Toon, J., Macandrew, H., McArdle, I., et al. (2017). ‘Let’s Talk – Interoperability between University CRIS/IR and Researchfi sh: A Case Study from the UK’, Procedia Computer Science, 106: 220–31. DOI: 10.1016/j.procs.2017.03.019
Costas, R., & van Leeuwen, T. N. (2012). ‘Approaching the “reward triangle”: General analysis of the presence of funding acknowledgments and “peer interactive communication” in scientific publications’, Journal of the American Society for Information Science and T echnology, 63/8: 1647–61. DOI: 10.1002/asi.22692
Costas, R., & Yegros-Yegros, A. (2013). ‘Possibilities of funding acknowledgement analysis for the bibliometric study of research funding organizations: Case study of the Austrian Science Fund (FWF)’. Gorraiz J., Schiebel E., Gumpenberger C., Hörlesberger M., & M oed H. (eds) Proceedings of ISSI 2013 – 14th International Society of Scientometrics and Informetrics Conference, Vol. 2, pp. 1401–8. Vienna: Austrian Institute of Technology.
Crotty, D. (2024). ‘Variability, Irregular Publisher Metadata, and the Ongoing Evolution of Databases Complicates Reproducibility in Bibliometrics Research’. The Scholarly Kitchen. Retrieved from <https://scholarlykitchen.sspnet.org/2024/08/15/variability-bad-publisher-metadata-and-the-ongoing-evolution-of-databases-makes-bibliometrics-research-reproducibility-difficult/>
Culbert, J. H., Hobert, A., Jahn, N., Haupka, N., Schmidt, M., Donner, P., & Mayr, P. (2025). ‘Reference coverage analysis of OpenAlex compared to Web of Science and Scopus’, Scientometrics, 130/4: 2475–92. DOI: 10.1007/s11192-025-05293-3
Dappert, A., Farquhar, A., Kotarski, R., & Hewlett, K. (2017). ‘Connecting the Persistent Identifier Ecosystem: Building the Technical and Human Infrastructure for Open Research’, Data Science Journal, 16. DOI: 10.5334/dsj-2017-028
Delgado-Quirós, L., & Ortega, J. L. (2024). ‘Completeness degree of publication metadata in eight free-access scholarly databases’, Quantitative Science Studies, 5/1: 31–49. DOI: 10.1162/qss_a_00286
Demes, K. (2026). ‘Funding metadata in OpenAlex’. OpenAlex blog. Retrieved March 1, 2026, from <https://blog.openalex.org/funding-metadata-in-openalex/>
Ebadi, A., & Schiffauerova, A. (2016). ‘How to boost scientific production? A statistical anal ysis of research funding and other influencing factors’, Scientometrics, 106/3: 1093–11, DOI: 10.1007/s11192-015-1825-x
van Eck, N. J., & Waltman, L. (2025). ‘Crossref as a source of open bibliographic metadata’. DOI: 10.31222/osf.io/smxe5_v2
El-Ouahi, J. (2024). ‘Research funding in the Middle East and North Africa: analyses of ackn owledgments in scientific publications indexed in the Web of Science (2008–2021)’, Sci entometrics, 129/6: 2933–68. DOI: 10.1007/s11192-024-04983-8
Eum, S., Mazoni, A. F., Seo, J. H., Park, J., & Costas, R. (2025). ‘Investigating the Metadata of Local Publications and the Tension Between Local and Global Object Identifiers: The Case of the Korea Citation Index (KCI)’. DOI: 10.31235/osf.io/6ch2v_v1
Fear, K., & Donaldson, D. R. (2012). ‘Provenance and credibility in scientific data repositorie s’, Archival Science, 12/3: 319–39. DOI: 10.1007/s10502-012-9172-7
Frank, R. D. (2022). ‘Risk in trustworthy digital repository audit and certification’, Archival S cience, 22/1: 43–73. DOI: 10.1007/s10502-021-09366-z
Gibson, D., van Honk, J., & Calero-Medina, C. (2022). ‘Acknowledging the Difficulties: A Case Study of a Funding Text’. Leiden Madtrics.
Grassano, N., Rotolo, D., Hutton, J., Lang, F., & Hopkins, M. M. (2017). ‘Funding Data from Publication Acknowledgments: Coverage, Uses, and Limitations’, Journal of the Associ ation for Information Science and Technology, 68/4: 999–1017. DOI: 10.1002/asi.23737
Greenberg, J. (2017). ‘Big Metadata, Smart Metadata, and Metadata Capital: Toward Greater Synergy Between Data Science and Metadata’, Journal of Data and Information Scienc e, 2/3: 19–36. DOI: 10.1515/jdis-2017-0012
Haupka, N., Culbert, J. H., Schniedermann, A., Jahn, N., & Mayr, P. (2025). ‘Analysis of the Publication and Document Types in OpenAlex, Web of Science, Scopus, PubMed and S emantic Scholar’, Quantitative Science Studies, 1–22. DOI: 10.1162/QSS.a.406
Hodgson, D., Fred, M., Bailey, S., & Hall, P. (Eds). (2019). The Projectification of the Public Sector. New York : Routledge, 2019. | Series: Routledge critical studies in public management: Routledge. DOI: 10.4324/9781315098586
van Honk, J., Calero-Medina, C., & Costas, R. (2016). ‘Funding Acknowledgements in the W eb of Science: inconsistencies in data collection and standardization of funding organizat ions’. Proceedings of the 21st International Conference on Science and Technology Indicators – STI 2016, pp. 90–6.
Ihli, M. (2017). ‘Methods for Synthesis of Funding Agency & Publisher Data’. Proceedings o f the 6th International Workshop on Mining Scientific Publications, pp. 57–63. New Yor k, NY, USA: ACM. DOI: 10.1145/3127526.3127537
International Science Council. (2021). Opening the record of science: making scholarly publi shing work for science in the digital era. Paris. Retrieved from <https://council.science/? post_type=publications&p=28587>. DOI: 10.24948/2021.01
Jang, H. (2025). ‘Impact of funding size on research outputs: evidence from South Korea’s government-funded programme’, Asian Journal of Technology Innovation, 33/1: 23–44. DOI: 10.1080/19761597.2024.2317192
de Jonge, H. (2025). ‘Celebrating one year of Crossref Grant IDs at NWO’. DOI: 10.64000/dvqke-j4v69
Kim, D., & Hwang, J. (2022). ‘Is renewable energy more favorable to diversity than conventi onal energy sources on R&D performance?’, Science and Public Policy, 49/4: 646–58. DOI: 10.1093/scipol/scac016
Kim, M., Han, Y. S., & Cho, R. M. (2025). ‘The role of cooperative R&D in innovation perfo rmance of SMEs: evidence from South Korean materials, parts, and equipment firms’, Asian Journal of Technology Innovation, 33/1: 126–51. DOI: 10.1080/19761597.2024.2349210
Koier, E., & Horlings, E. (2015). ‘How accurately does output reflect the nature and design of transdisciplinary research programmes?’, Research Evaluation, 24/1: 37–50. DOI: 10.1093/reseval/rvu027
Kokol, P. (2023). ‘Discrepancies among Scopus and Web of Science, coverage of funding inf ormation in medical journal articles: a follow-up study’, Journal of the Medical Library Association, 111/3: 703–9. DOI: 10.5195/jmla.2023.1513
Kokol, P., & Blažun Vošner, H. (2018). ‘Discrepancies among Scopus, Web of Science, and PubMed coverage of funding information in medical journal articles’, Journal of the Me dical Library Association, 106/1. DOI: 10.5195/jmla.2018.181
Korea Institute of Science and Technology Evaluation and Planning. (2017). Report on the P erformance Analysis of National Research and Development Programs in 2015. (In Korean: 한국과학기술기획평가원. (2017). 2015년도 국가연구개발사업 성과분석 보고서.)
——. (2023). ‘Implementation Plan for Survey and Analysis of National R&D Programs in 2 023 and Manual for Input’. (In Korean: 한국과학기술기획평가원. (2023). 2023년도국가연구개발사업 조사·분석 실시계획과 입력 매뉴얼.)
Korea Research Foundation. (2009). ‘2008 management guidelines for academic research pro jects’. (In Korean: 한국학술진흥재단. (2009). 2008 학술연구과제관리규칙.)
Korean Government Ministries Concerned. (2005). ‘Plan for the establishment of a comprehe nsive national science and technology information system’. National Science and Techn ology Council. (In Korean: 관계부처 합동. (2005). 국가과학기술종합정보시스템 구 축계획. 국가과학기술위원회.)
Kramer, B., & de Jonge, H. (2022). ‘The availability and completeness of open funder metad ata: Case study for publications funded by the Dutch Research Council’, Quantitative Science Studies, 3/3: 583–99. DOI: 10.1162/qss_a_00210
Kwon, H. Y., & Kwon, I. (2019). ‘R&D Spillovers for Public R&D Productivity’, Global Economic Review, 48/3: 334–49. DOI: 10.1080/1226508X.2019.1638812
Lee, H., Baek, S., Shin, J., & Kim, H. (2020). A study on policies for declaring acknowledgments in research articles. National Research Foundation of Korea. (In Korean: 이효빈 외. (2020). 연구논문의 사사(謝辭) 표기 정책에 관한 연구. 한국연구재단.)
Lee, J., Shin, K., Kim, H., & Hwang, J. (2024). ‘Efficiency of Innovation Policy with Different Types of R&D Planning: Evidence from South Korea’s Information and Communicat ion Technology Sector’, Journal of the Knowledge Economy, 16/1: 630–62. DOI: 10.1007/s13132-024-01947-4
Lee, K., Choi, S., & Yang, J.-S. (2021). ‘Can expensive research equipment boost research an d development performances?’, Scientometrics, 126/9: 7715–42. DOI: 10.1007/s11192-021-04088-6
Lin, D., Crabtree, J., Dillo, I., Downs, R. R., Edmunds, R., Giaretta, D., De Giusti, M., et al. (2020). ‘The TRUST Principles for digital repositories’, Scientific Data, 7/1: 144. DOI: 10.1038/s41597-020-0486-7
Liu, W. (2020). ‘Accuracy of funding information in Scopus: a comparative case study’, Scientometrics, 124/1: 803–11. DOI: 10.1007/s11192-020-03458-w
MacLean, M., Davies, C., Lewison, G., & Anderson, J. (1998). ‘Evaluating the research activ ity and impact of funding agencies’, Research Evaluation, 7/1: 7–16. DOI: 10.1093/rev/7.1.7
Meddings, K. (2013). ‘FundRef: connecting research funding to published outcomes’, Insight s: the UKSG journal, 26/3: 272–6. DOI: 10.1629/2048-7754.98
Ministry of Education, Science and Technology. (2009). ‘Administration guidelines for the su pport program for academic research in the humanities and social sciences’. (In Korean: 교육과학기술부. (2009). 인문사회분야 학술연구지원사업 처리규정, 교육과학기 술부훈령 제111호.)
Ministry of Science and ICT. (2019). ‘National Science and Technology Information Service (NTIS) 5.0 Basic Plan (2019-2021)’. Presidential Advisory Council on Science and Tec hnology. (In Korean: 과학기술정보통신부. (2019). 국가과학기술지식정보서비스(N TIS) 5.0 기본계획(2019~2021).)
Mongeon, P., Bowman, T. D., & Costas, R. (2023). ‘An open data set of scholars on Twitter’, Quantitative Science Studies, 4/2: 314–24. DOI: 10.1162/qss_a_00250
Morillo, F., & Álvarez-Bornstein, B. (2018). ‘How to automatically identify major research s ponsors selecting keywords from the WoS Funding Agency field’, Scientometrics, 117/ 3: 1755–70. DOI: 10.1007/s11192-018-2947-8
Mugabushaka, A.-M. (2021). ‘Linking Publications to Funding at Project Level: A curated da taset of publications reported by FP7 projects’,.
Mugabushaka, A.-M., Baglioni, M., Bardi, A., & Manghi, P. (2021). ‘Scholarly outputs of E U Research Funding Programs: Understanding differences between datasets of publicati ons reported by grant holders and OpenAIRE Research Graph in H2020’,.
Mugabushaka, A.-M., van Eck, N. J., & Waltman, L. (2022). ‘Funding COVID-19 research: I nsights from an exploratory analysis using open data infrastructures’, Quantitative Science Studies, 3/3: 560–82. DOI: 10.1162/qss_a_00212
Paul-Hus, A., Díaz-Faes, A. A., Sainte-Marie, M., Desrochers, N., Costas, R., & Larivière, V. (2017). ‘Beyond funding: Acknowledgement patterns in biomedical, natural and social sciences’, (L. Bornmann, Ed.)PLOS ONE, 12/10: e0185578. DOI: 10.1371/journal.pone. 0185578
Pipino, L. L., Lee, Y. W., & Wang, R. Y. (2002). ‘Data quality assessment’, Communications of the ACM, 45/4: 211–8. DOI: 10.1145/505248.506010
Powell, K. (2019). ‘Searching by grant number: comparison of funding acknowledgments in NIH RePORTER, PubMed, and Web of Science’, Journal of the Medical Library Association, 107/2: 172–8. DOI: 10.5195/jmla.2019.554
Reiche, K. J., & Hofig, E. (2013). ‘Implementation of Metadata Quality Metrics and Applicat ion on Public Government Data’. 2013 IEEE 37th Annual Computer Software and Appli cations Conference Workshops, pp. 236–41. IEEE. DOI: 10.1109/COMPSACW.2013.32
Rigby, J. (2011). ‘Systematic grant and funding body acknowledgement data for publications: new dimensions and new controversies for research policy and evaluation’, Research E valuation, 20/5: 365–75. DOI: 10.3152/095820211X13164389670392
Rigby, J., & Julian, K. (2014). ‘On the horns of a dilemma: does more funding for research le ad to more research or a waste of resources that calls for optimization of researcher portf olios? An analysis using funding acknowledgement data’, Scientometrics, 101/2: 1067–7, DOI: 10.1007/s11192-014-1259-x
Schares, E. (2024). ‘Comparing Funder Metadata in OpenAlex and Dimensions’, OpenISU. DOI: 10.31274/b8136f97.ccc3dae4
Scheidsteger, T., & Haunschild, R. (2022). ‘Comparison of metadata with relevance for bibli ometrics between Microsoft Academic Graph and OpenAlex until 2020’,. DOI: 10.3145/epi.2023.mar.09
Scheidsteger, T., Haunschild, R., Hug, S., & Bornmann, L. (2018). ‘The concordance of field-normalized scores based on Web of Science and Microsoft Academic data: A case study in computer sciences’,.
Sirtes, D. (2013). ‘Funding acknowledgements for the German Research Foundation (DFG). The dirty data of the Web of Science database and how to clean it up’. Gorraiz J., Schieb el E., Gumpenberger C., Hörlesberger M., & Moed H. (eds) Proceedings of ISSI 2013 – 14th International Society of Scientometrics and Informetrics Conference, Vol. 1, pp. 78 4–95. Vienna: Austrian Institute of Technology.
Thelwall, M., Simrick, S., Viney, I., & Van den Besselaar, P. (2023). ‘What is research funding, how does it influence research, and how is it recorded? Key dimensions of variatio n’, Scientometrics, 128/11: 6085–106. DOI: 10.1007/s11192-023-04836-w
Visser, M., van Eck, N. J., & Waltman, L. (2021). ‘Large-scale comparison of bibliographic d ata sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic’, Quantitative Science Studies, 2/1: 20–41. DOI: 10.1162/qss_a_00112
Zhang, L., Cao, Z., Shang, Y., Sivertsen, G., & Huang, Y. (2024). ‘Missing institutions in Op enAlex: possible reasons, implications, and solutions’, Scientometrics, 129/10: 5869–91. DOI: 10.1007/s11192-023-04923-y
Appendix A. Matching procedure: technical details
A.1 Attribute-specific scoring
Matching scores for each attribute were assigned as follows.
Title: The score is based on title similarity calculated using the Levenshtein distance, according to the following formula, where 𝑡𝑋 denotes the title of record 𝑋, 𝐷(𝑎, 𝑏) equals the Levenshtein distance between string 𝑎 and 𝑏, and 𝐿(𝑋) denotes the length of string 𝑋.
𝑚𝑡𝑖𝑡𝑙𝑒 = 1 − 𝐷(𝑡𝑎, 𝑡𝑏)/ max (𝐿(𝑡𝑎), 𝐿(𝑡𝑏))
Source: The score is similarly based on the similarity of source titles, calculated using the Levenshtein distance, according to the following formula, where 𝑠𝑋 denotes the source title of record 𝑋. When multiple title variants exist for a single source, the variant yielding the highest matching score is selected. If the ISSNs of two records match, a score of 1 is assigned.
𝑚𝑠𝑜𝑢𝑟𝑐𝑒 = 1 − [𝐷(𝑠𝑎, 𝑠𝑏) − |𝐿(𝑠𝑎) − 𝐿(𝑠𝑏) |]/min (𝐿(𝑠𝑎), 𝐿(𝑠𝑏))
Publication Year: A score of 1 is assigned if the publication years are identical; 0.5 when the difference is one year; 0.25 for a two-year difference; and 0 when the difference exceeds two years. This accommodates cases in which researchers report early-access or acceptance dates, as well as discrepancies between online and print publication dates.
Volume, Issue, Beginning Page, and Article Number: For each attribute, a score of 1 is assigned if it matches exactly, and 0 otherwise.
A.2 Weight and thresholds
Different attribute-specific weights and thresholds were applied at each step, as shown in Table A1. For example, the matching score between NTIS and WoS in Step 1 is given by the following formula, where 𝑚𝑋 denotes the matching score of each attribute 𝑋. A pair is retained if the resulting score meets or exceeds the threshold of 60 in this step.
𝑆𝑐𝑜𝑟𝑒𝑤𝑜𝑠,𝑠𝑡𝑒𝑝1 = 40𝑚𝑡𝑖𝑡𝑙𝑒 + 20𝑚𝑠𝑜𝑢𝑟𝑐𝑒 + 10𝑚𝑦𝑒𝑎𝑟 + 5𝑚𝑣𝑜𝑙𝑢𝑚𝑒 + 3𝑚𝑖𝑠𝑠𝑢𝑒 + 1𝑚𝑏𝑒𝑔𝑖𝑛𝑛𝑖𝑛𝑔_𝑝𝑎𝑔𝑒 + 1𝑚𝑎𝑟𝑡𝑖𝑐𝑙𝑒_𝑛𝑢𝑚𝑏𝑒𝑟
DOI was excluded from the matching score calculation because it is not available for all NTIS records; including it would have disproportionately penalised records lacking DOIs. Source similarity was excluded from score calculations in Steps 2 and 3, as candidate documents were already restricted to those sharing the same ISSN or journal title, meaning the source similarity metric would assign a score of 1 to all such pairs and thus contribute no meaningful discrimination. Multiple combinations of weights and thresholds were tested, and the configuration yielding the highest precision and recall was adopted.
Table A1. Weights and thresholds for matching attributes at each step
|
Step 1:
DOI-based |
Step 2:
ISSN-based |
Step 3:
Journal name-based |
Step 4:
All match keys |
| WoS |
OpenAlex |
WoS |
OpenAlex |
WoS |
OpenAlex |
WoS |
OpenAlex |
| Title |
40 |
40 |
40 |
40 |
40 |
40 |
35 |
35 |
| Source |
20 |
20 |
|
|
|
|
25 |
25 |
| Year |
10 |
10 |
18 |
18 |
14 |
14 |
10 |
10 |
| Volume |
5 |
5 |
10 |
10 |
12 |
12 |
5 |
5 |
| Issue |
3 |
3 |
6 |
7 |
6 |
8 |
3 |
3 |
| Beginning
page |
1 |
2 |
3 |
5 |
4 |
6 |
1 |
2 |
| Article
number |
1 |
|
3 |
|
4 |
|
1 |
|
| Threshold
score |
60 |
60 |
50 |
50 |
50 |
50 |
60 |
60 |
Appendix B. Reproduction of national science and technology statistics
As an additional validation of the matching results, I reproduced national science and technology statistics by aggregating funded publications using the contribution rate field in NTIS records. The number of funded publications is counted annually to assess national R&D performance and to compile official statistics. Accordingly, if the counting method was accurately replicated and NTIS records were correctly matched to documents indexed in bibliographic databases, consistency with national statistics provides further evidence of the reliability of the matching results.
The reproduction exercise also aimed to clarify the process by which national statistics are compiled. Because NTIS records do not correspond one-to-one with actual publications (each record does not represent a unique document) and because document-specific persistent identifiers such as DOIs are incompletely recorded, simply counting NTIS records does not yield an accurate number of funded publications. Although KISTEP manually identifies and aggregates funded publications, this process was neither transparent nor documented, preventing others from replicating the process and arriving at the same results. In this context, reproducing national statistics not only serves as a validation of the matching results but also provides a replicable protocol for analysing funded publication counts in settings where unique identifiers are incomplete or unavailable.
By aggregating contribution rates across records, I closely approximated the official statistics on the number of funded publications. Moreover, the number of distinct documents identified in WoS and OpenAlex closely aligned with the reproduced national counts, as shown in Figure A1. The discrepancies between national statistics and the WoS- and OpenAlex-based counts were 0.9% and 1.5%, respectively, largely attributable to NTIS records that could not be matched to corresponding bibliographic documents. On this basis, I conclude that aggregation of funded publications in NTIS and document-based counting using matched records from WoS and OpenAlex produce highly consistent results.

Figure B1. Reproduction of national statistics for funded publications
/