Published at MetaROR

November 25, 2024

Table of contents

Available versions
Cite this article as:

Porter, S. J., & Hook, D. W. (2024). The Rise and Fall of the Initial Era. arXiv preprint arXiv:2404.06500.

The rise and fall of the initial era

Simon Porter1 Email, Daniel Hook1 Email

1 Digital Science

Originally published on May 8, 2024 at: 

Abstract

Bibliographic data is a rich source of information that goes beyond the use cases of location and citation -- it also encodes both cultural and technological context. For most of its existence, the scholarly record has changed slowly and hence provides an opportunity to gain insight through its reflection of the cultural norms of the research community over the last four centuries. While it is often difficult to distinguish the originating driver of change, it is still valuable to consider the motivating influences that have led to changes in the structure of the scholarly record. An "initial era" is identified during which initials were used in preference to full names by authors on scholarly communications. Causes of the emergence and demise of this era are considered as well as the implications of this era on research culture and practice.

Introduction

In the contemporary discourse on technology, dominated by references to digital and electronic innovations, the notion of the research article as a form of technology may appear incongruous. However, an exploration of the research article through the lens of technology is critical for a comprehensive understanding of its role, interactions with other technological forms, and its consequent impact on society. The research article, viewed technologically, is a significant construct, with a long-standing history of shaping social norms and establishing institutions that extend their influence across the research community, irrespective of disciplinary boundaries, geographical locations, or historical periods [1]. This technology has achieved ubiquity on three distinct levels: spatial, with researchers globally engaging with and understanding research articles under shared assumptions; temporal, allowing for the contextual comprehension of older articles through a slow evolution of the format; and disciplinary, with a cross-disciplinary recognition of the rigorous scrutiny and scientific methodology underpinning the work. These characteristics are critical for research to function, as an incremental activity that builds on prior results and knowledge.

The intrinsic characteristics of research papers have rendered them the foundational elements of research communication and, crucially, the conduits of trust among researchers, transcending spatial, temporal, and disciplinary divides. This established trust facilitates incremental research and underpins the development and cohesion of the global research community. Analogous to economic institutions, the norms surrounding research papers enable researchers to make assumptions similar to the reliability of contracts in legally robust countries, thus enabling international academic transactions. Beyond facilitating trust, the structured format of a research paper—detailing the research’s specifics, the researchers, the location and timing of the research, funding sources, and relevant previous work—supports the provenance and contextualisation essential for the credibility of its communicated results.

The interaction between technology and its consequent influence over its users and communities is a well-documented phenomenon; however, possibly due to the long-lived and slow-changing nature of its underlying format, the research paper stands out for its persistence over the centuries. Over approximately the first 300 years of formalised research communication, dating back to the 1660s, the pace of change in research publication formats has been gradual. Only in the last half-century has the rapid transformation of research practices necessitated a quicker evolution of this technology.

At its core, the research paper remains a rare example of a 17th-Century technology in current use. It is a technology that originally emerged from a very different time to serve a very different context. It was the Age of Empire, where research communications were characterised by brief correspondences, often containing hand-drawn diagrams or concise result tables, exchanged among a small, affluent elite. Originating in the club culture of coffee houses in cities like London and Paris, this technology now underpins a vast, global research ecosystem, supporting millions of contributors and a burgeoning diversity of cultures, geographies, and subjects [234].

The study of the representation of author names provides us with an insight into the structure of research culture and the sociology of research. This has been well-studied with the implications of gender bias in various aspects of the scholarly record being examined in detail [56789]. It is not the focus of this article to rehash these arguments but rather to add to the data supporting these arguments. In a period where there are high proportions of researchers using full-form names, we are able potentially to determine more about the gender and background of the participants in the research ecosystem and thus learn about the geographical and discipline-based diversity of the participants in the research ecosystem. When we enter periods where initial-form is used, data become less rich and we are able to determine less.

This study leverages the slow evolution of the research paper to conduct a form of “bibliometric archaeology” [10] over the digital research record [1112], examining the interplay between technological and societal changes as encapsulated in the history of the research paper’s format. By focusing on seemingly minor details, such as the presentation of author names, from full given names (“full form”) to initials only (“initial form”), this paper seeks to uncover broader narratives, correlating shifts in norms with technological, geographic, disciplinary, and societal factors.

The paper is organised as follows: In the remainder of the current section we set a historical backdrop that gives an insight into the consistency of the data that we study. In the methodology section, we give an overview of the data sources and technology used with exemplar code, and we describe the approach that we have taken to classify authors and papers into the different cohorts required for the analyses described above. In the results section we present several analyses and interpret these. Finally, we conclude the paper with a brief discussion of the results and suggest future directions of research.

A brief history of metadata norms

The prevailing norm within contemporary academic discourse mandates the association of scholarly outputs—be it journal articles, conference proceedings, books, or other forms—with one or more authors. This practice, however, was not always a staple of scholarly communication. The explicit linkage of authors to their scholarly contributions, particularly when introducing novel results or claims, has evolved into a cornerstone of the modern research paradigm. This evolution is driven largely by the critical role of provenance in establishing the reliability of claims, facilitating subsequent research built upon these foundations, and the necessity for researchers to be credited with their findings as a means of securing funding and institutional support.

Despite the ubiquity of this norm, contemporary scholarly practices do witness exceptions, particularly in types of outputs that do not introduce novel findings, such as editorial articles, where the omission of author names remains more common.

In the nascent stages of formal scholarly communication, particularly within the “club” culture that characterised the early operations of the Philosophical Transactions of the Royal Society—regarded as the first formal research journal—the practice of attributing author names to scholarly works was not deemed essential. An analysis of the journal’s publications from its initial five years (1665-1669) reveals that only 171 out of 357 articles listed in the Dimensions database from Digital Science feature identifiable authors. Furthermore, during this period, not a single contribution to the journal provided the address or institutional affiliation of the author, underscoring a significant departure from contemporary standards of authorship and attribution.

Figure 1: A page from the Philosophical Transactions of the Royal Society from December 1669 demonstrating at once the similarity to a modern research article and the significant differences in the level and detail of metadata present in the article [13].

In the 17t⁢h Century, at the advent of formalised scholarly communication, there was a relatively small community and communications to the journal either arrived directly from a member of a scholarly society or were submitted via a member. Thus, names were often either considered immaterial, would have been known from context and are now lost to time, or were elegantly, and seemingly purposefully not mentioned as in the example from 1669 shown in Fig. 1, which reads

“A Letter Written by an Intelligent and Worthy English Man from Paris, to a Considerable Member of the R. Society in London, concerning some Transactions there, relating to the Experiment of the Transfusion of Blood.”

The letter is dated but not signed and the author is not explicitly named [13].

The genesis of any novel form of discourse inherently involves a period of adjustment, wherein norms and social contracts gradually emerge and solidify. This evolutionary process has led to a myriad of reasons for the inclusion or exclusion of specific details in formal communications, some of which mirror broader societal norms. Reflecting on the origins of scholarly communication, primarily rooted in British and European traditions, one can draw parallels with the literary conventions of the time. For instance, in Jane Austen’s ”Pride and Prejudice” (published in 1797), the first names of Mr. and Mrs. Bennet are never revealed. Similarly, Arthur Conan Doyle’s iconic characters, Holmes and Watson, introduced in 1887, seldom use first names in their interactions, reflecting the written norms of their times.

The reasons for omitting author identities can also stem from more problematic motives, such as the desire to conceal an author’s gender. Now famously, the Brontë sisters initially published under male pseudonyms as writing was not an appropriate activity for women in that period. More recently J.K.Rowling opted to use her initials to conceal her gender for commercial reasons, anticipating potential bias in her readership, before she became too well known and successful to be able to conceal her identity [14].

Anonymity or pseudonymity can be a choice not only to protect from persecution, it can also serve to protect authors from repercussions related to the voicing of controversial opinions or work that is not in step with the political regime where it is being carried out. The revolutionary ideas shared on the future of currency are an example in the seminal white paper on Bitcoin by Satoshi Nakamoto, where the identity, gender, nationality, or even the individuality or plurality of the author(s) remains unknown [15].

The evolution of the research article as technology

The role of technology in facilitating or restricting the disclosure of author information has pivotal import. For example, while current practices do not mandate the encoding of gender data in author identities, this omission has been significant in fields where research outcomes vary with the researcher’s gender [161718]. The decision not to maintain detailed records of authors’ gender or ethnicity is influenced by contemporary societal norms, yet future scholars might view this lack of information as a critical oversight, akin to how the absence of affiliation details in early scholarly works is perceived today.

Perhaps more importantly than the needs of future scholars, the increase in transparency afforded by first author names plays a daily role in how we experience research. First names, in the ethnicities and genders that they suggest, provide an, albeit imperfect, high-level reflection of the diversity of experiences that are brought to research. It is just as important to see ourselves reflected in the outputs of the research careers that we choose to pursue, as the voices that represent us on panels at conferences. This highlights the non-neutral role of technology in the representation of individual-related information.

The interplay between technological and societal changes, each influencing the other in complex ways, is also evident in the shifting patterns of co-authorship in scholarly papers. Historical analysis (see Fig. 2) reveals a shift in the modal average number of authors from one to gradually increasing numbers of participants over time, reflecting the changing sociology of research. Analysis of the locations of researchers shows the fundamental nature of research community [19].

The slow and steady evolution shown in Fig. 2 in collaboration may also be considered to be a societal parallel for the slow evolution in the nature of the technologies that underpin scholarly communication. The increase in authorship of a paper from a single author (modal average) from the genesis of scholarly communication until around 1960 followed by a shift, relatively rapidly, to five co-authors per paper as a modal average in the present day, gives us a sense of when technological and social changes started to facilitate and change norms. A more detailed analysis of evolution of co-authorship and its importance in addressing more complex problems is dealt with by Thelwall and Maflahi [20].

Figure 2: Development of averages of co-authorship numbers on academic output since 1665. The red line shows the modal number of co-authors, the yellow line shows the median number of co-authors, the blue line shows the mean average number of co-authors. The grey shaded area shows the standard deviation from the mean cutoff by the zero axis.

The 20th century witnessed a transformative shift in research practices, fundamentally driven by technological advancements. This period marked the emergence of “Big Science”, a paradigm characterised by large-scale scientific endeavours that required extensive collaboration across disciplines and often, across national borders. The sociological landscape of certain research fields underwent significant changes due to this shift, influencing and reshaping established norms within the scientific community. The impact of Big Science is particularly evident in the evolving patterns of authorship, as collaborations expanded to include hundreds or even thousands of contributors.

Illustrative of this trend, the growth in the number of research papers featuring extensive co-authorship is striking (see Fig. 3). Papers with more than 100 authors, indicative of large collaborative projects, began to appear with greater frequency, reflecting the complex, interdisciplinary nature of modern scientific inquiries (yellow region) starts in earnest in the late 1970s 1. This trend is further accentuated in research requiring substantial resources, such as particle physics experiments, where the collective expertise and effort of hundreds of scientists are essential to the project’s success. Papers with over 500 co-authors (blue) and, notably, more than 1000 co-authors (red) have become increasingly common, underscoring the scale of collaboration in certain areas of research.

Figure 3: Number of papers in each year with more than 100 (yellow area), 500 (blue area) and 1000 (red area) co-authors respectively.

From the overarching observations and detailed examples discussed, three key insights emerge that elucidate the evolution of scholarly communication:

  1. 1. Gradual Evolution of the Scholarly Record: The development of conventions around the scholarly record, such as the norm of naming authors on papers, reveals a slow evolutionary process. It highlights that it took nearly two centuries to establish a general expectation for author attribution on scholarly works. This gradual change underscores the inertia inherent in academic traditions and the time it takes for new norms to solidify across the research community.

  2. 2. Accelerated Changes in Recent Times: The pace at which the scholarly record has been transforming has markedly increased in recent years. This acceleration is evident in the rapid shift from a median of two authors to five authors per paper across a span of 60 years, a change driven by the increasing frequency, volume, and data size of scholarly outputs. This contrasts sharply with the nearly 300-year period it previously took for such a demographic shift in authorship, highlighting how technological advancements and the expanding scope of collaborative research have expedited changes in scholarly communication practices.

  3. 3. Reflective Richness of Data: The data surrounding scholarly communication not only mirrors the evolution of the research community but has, in recent times, also become indicative of the technologies that facilitate this communication. The growing complexity and interconnectedness of research outputs, evidenced by the increasing number of co-authors and the expansion of collaborative networks, reflect broader shifts in the technologies available for research and dissemination. This richness in data offers profound insights into the dynamics of scholarly practices and their evolution over time.

On a practical level, the significance of author names extends beyond mere attribution. In the context of evaluation for promotions, tenure, grant assessments, and many other forms academic recognition, the identity of authors, corroborated by associated identifiers, plays a crucial role. Indeed, research is becoming increasingly centred around quantification [22]. The trust placed in a particular paper, and by extension, in communities of researchers, is increasingly dependent on the clarity and reliability of authorship information. Furthermore, this information is pivotal in selecting referees, assessing potential conflicts of interest, and facilitating a myriad of other critical academic processes. The evolution of scholarly communication, therefore, not only reflects changes in the academic landscape but also emphasises the growing importance of clear and reliable attribution in maintaining the integrity and trustworthiness of scholarly work.

Methodology

In the study presented here, we explore the change in the representation of author names in scholarly research articles. As established in the Introduction, we define the terms initial form to refer to an author name where only the initials are supplied and no given name is supplied, and full form to refer an author name that includes at least one fully stated given name. We use Dimensions from Digital Science as the data source for our exploration. At the time of writing Dimensions contains over 144 million publications, of which approximately 112 million are journal articles. Much of the analysis shown in this paper is based on queries that can be run directly on the Dimensions database in the Google BigQuery environment without the need for further coding in, for example, a Python environment and hence are accessible to a broad audience [1123].

Although we have included several graphs and examples regarding scholarly output before 1945 we recognise that publications volumes are smaller further back in time. Thus, we define two epochs in publication, one before 1945 which we refer to as the anecdotal epoch and the other after 1945 that we call the statistical epoch. It is noteworthy that Dimensions, our primary data source, is limited in the validity of its data for earlier times in the anecdotal epoch. The data source was not constructed with the current use case in mind and hence coverage of content in the 17th, 18th and even 19th centuries is reliant on publisher and community interest in awarding persistent identifiers to research outputs of these periods in order for them to be included in Dimensions. In addition, mappings to institutions that have disappeared or changed their name over the earlier part of the last 350 years may not be fully represented in geographical analyses – similarly, country references will be to modern countries not countries that have existed in the past – we take no account of moving boundaries, merely imposing the current world view back in time for simplicity of analysis.

Broadly speaking we have set out not to use statistical methods or to take inferences based on data in the anecdotal epoch – such figures that depict these data are intended to be illustrative, sometimes highlighting the limitations of Dimensions for the period, and at other times to add anecdotal colour and context to our discussions and arguments. Data concerning the statistical epoch, is more trustworthy both from a structural perspective (systems existed to more completely capture the scholarly record and it is a period during which there remains significant academic interest in reading and making reference to the material, thus the literature of this period has much better coverage of persistent identifiers) as well as from a volume perspective. Fig. 2 shows the growth in volume of publications (using the restrictions specified in Listing 1).

In terms of both the quality and volume of the data, the transition from the anecdotal to the statistical epoch is not sharply defined but happens gradually. However, our choice of 1945 is not completely arbitrary. The current authors have previously noted that 1945 was the year in which the centre of mass of global research stopped its march toward North America and began to move back toward Europe [11]. It was also the year in which Vannevar Bush published the Endless Frontier [24] and, as such, is a good point in time from which to date the modern approach to research. Indeed, both Figs 3 and 4 show a marked change in behaviour between 1940 and 1950.

On a technical note, during the Dimensions data ingestion process author details are gathered from across the available sources that Dimensions draws on. Each author has a last name and a first name, a unique identifier assigned within Dimensions (where it is possible to determine), a list of institutional affiliations mapped to unique identifiers (GRIDs, where resolvable), an ORCID (again, where resolvable) that have been asserted in relation to a given output, and a marker of whether the author is listed as a corresponding author. For the majority of publishers this is delivered via a JATS XML feed, but smaller publishers or those without the technical capability to deliver JATS XML, Crossref data is used or, with the permission of the publisher, data is crawled directly from their website. Once cleaned and enhanced in the Dimensions data pipeline, the data are parsed into the Dimensions database and loaded into Google BigQuery.

Each analysis that we present takes either an author-centric or paper-centric view. That is to say we either examine the proportion of papers published in a year with certain characteristics or we examine the characteristics of authors that have published in the year directly. As a side note, we observe that the active publishing community is a reasonable proxy for the global research population but at each point in the timeline that we engage with, there are good reasons why this cannot be taken to be more than a representative sample. For example, in the early years of the scholarly record (circa 1665) publication was a new activity that was not engaged with by every academic and social structures of the time tended to exclude women. On the other hand, in the present day, a significant proportion of global research takes place in proprietary settings and hence is not published in the scholarly record. Thus, all the comments that we make in the paper need to be taken within these constraints in mind.

-- RAW table: Split the first name string into sections 

and prepare the first three for analysis, gather other 

key locational information that we might need later.

SELECT p.id,
p.year,
author.researcher_id,
ady.country_code,
ady.grid_id,
author.first_name,
author.last_name,
LENGTH(author.first_name) total_length,
LENGTH(SPLIT(author.first_name,’ ’)[SAFE_OFFSET(0)]) var1_one_length,
STRPOS(SPLIT(author.first_name,’ ’)[SAFE_OFFSET(0)],".") var1_one_dot,
LENGTH(SPLIT(author.first_name,’ ’)[SAFE_OFFSET(1)]) var1_two_length,
STRPOS(SPLIT(author.first_name,’ ’)[SAFE_OFFSET(1)],".") var1_two_dot,
LENGTH(SPLIT(author.first_name,’.’)[SAFE_OFFSET(0)]) var2_one_length,
STRPOS(SPLIT(author.first_name,’.’)[SAFE_OFFSET(0)],".") var2_one_dot,
LENGTH(SPLIT(author.first_name,’.’)[SAFE_OFFSET(1)]) var2_two_length,
STRPOS(SPLIT(author.first_name,’.’)[SAFE_OFFSET(1)],".") var2_two_dot,
LENGTH(SPLIT(author.first_name,’’)[SAFE_OFFSET(2)]) var1_three_length,
LENGTH(SPLIT(author.first_name,’.’)[SAFE_OFFSET(2)]) var2_three_length,
FROM dimensions-ai.data_analytics.publications p,
UNNEST(authors) author
LEFT OUTER JOIN UNNEST(author.affiliations_address) ady
WHERE ARRAY_LENGTH(authors) < 50 and p.type=’article’

Listing 1: Core SQL Query for Dimensions on Google BigQuery that takes the name string and begins the process of allowing the classification of names into initial form or full form.

It would be a simple matter to test the first name to see if it is of length one or two characters (either an initial or an initial and a “.”). However, this approach would miss people with multiple initials or with a two character first name such as “Bo”. Hence, we use a slightly more complex approach were we examine the first three segmented tokens in the first name field in Dimensions. We test for the length of the field, the position of the character “.” . This allows us to correctly classify edge cases such as authors that are known principally, and hence who publish by, their second name such as J. Robert Oppenheimer (see, for example [25]).

A further complexity is that, in many cases, names are preserved in both their original and transliterated forms in Dimensions. In these cases a whole name is a single character, for example in Kanji or Katakana (see, for example [26]) and hence a whole name can be mistaken for an initial. Fortunately, this problem is a small one in the context of the current work since, as of the date of writing, only 250,991 of 8,456,140 publications in Chinese (or around 3%) have the potential to contain Kanji characters, that might have have been transliterated into English. There is a potentially bigger issue for Japan in that 1,152,602 of 4,460,000 Japanese language publications (or around 25%) have the potential to contain Kanji or Katakana characters without transliteration. However, deeper examination suggests that just 176,177 of 2,741,338 authors (or around 6.5%) with a Japanese address have a single character first name – which may legitimately be a transliterated first initial or a Japanese character that could be misclassified as an initial. While this number is not insignificant, we only wish to pick out high-level trends in the data in this article and hence this is negligible. We have not commented on how the percentage of Chinese language and Japanese language publications requiring translation has changed with time, however, this point is currently moot as global research has been dominated by countries with roman alphabets in the past and, again, for an analysis that requires only a high-level trend, it is reasonable to assume that the variance introduced by this effect does not change the overall appearance of the line.

We have limited the query in Listing 1 to include only outputs that are classed as papers, since books and conference proceedings have different qualities and characteristics that we don’t wish to cloud the current discussion. We have also removed papers with more than 50 co-authors as these papers also have different characteristics.

We have built further aggregations on top of the core table provided by Listing 1 to support the other analyses presented in this paper. Once the data are presented in the format from Listing 1 we perform aggregations that give a determination of the names of authors at both an author and a paper level. In the case of an author we classify their name as short form or long form and we classify a paper as initial form (for all co-authors), full form (for all co-authors), “Mixed” for mixtures of the prior two types or “Undetermined” when we can’t map names to either short or long form categories. The structure of Listing 1 allows us then to study how the data change with time, with subject category and with country.

In exploring the issue of gender we used genderize.io under licence. The genderize.io tool provides a dataset with probabilities that a given name is statistically associated with either men or women, or “undetermined” on both a nation-neutral and national basis. The tool is well-used in the study of gender in academia [2728]. This is a statistical analysis and hence cannot attempt to access the subtle and complex issues of gender identity, but rather it focuses on how gender is represented in the historical scholarly record. We have employed a methodology where we matched name variations in author names in Dimensions to the results obtained for that name variation in genderize.io. The results were stored in a private table in Google BigQuery. We then matched author names, where given, to data in this table. We used the current university of the academic to provide nationality information and, where data was either unavailable or indeterminate, we defaulted to using data without a national context. Obviously, this approach is open to errors in that, especially in more recent years, researchers are much more likely to travel during their careers and hence their current university may not be culturally relevant for such an analysis. Many names are also used much more internationally now and hence those that are traditionally nationally associated may not be used in context. It is also the case that some researchers will originate from countries where a Western name is adopted for business purposes. In this final case, the names are often more central to common usage as they are designed to mimic Western-style naming conventions and may even improve the data. The point of this paper is not, however, to do a detailed gender analysis and we believe that our approach is sufficiently robust to be serviceable in the current context.

Figure 4: Volume of scholarly publication using the restrictions specified in Listing 1 for comparability with other plots in the paper. This plot shows a clear cutoff around 1940 at which point there is sufficient data for the reasonable interpretation of statistical approaches.

Results

Paper-based analysis

As indicated in the introduction, the historical use of names in the scholarly record has been uneven and may, to a certain extent, reflect societal norms. While, in the early years of Dimensions’ coverage, the volumes of data are small, Fig. 5shows that until around 1800 it appeared to be normal to use one’s full name in scholarly correspondence (despite the fascinating cases discussed in the introduction). It is also important to acknowledge the lens through which we are examining the past is imperfect. The work of Lefanu [29] makes it evident that there was a lively and thriving research discussion prior to 1800, however, records have either not been preserved or not been transferred into digital realms, meaning that they are invisible to our study.

Figure 5: Proportion of papers in which all authors state their names using initials (“Initial form %” – red line), versus all authors stating their full names (“Full form %” – blue line) versus some authors stating their full name while others state their initials (“Mixed %” – green line. “Undetermined %” includes edge cases that we have not programmed for including no first name or no name at all. We note four distinct eras of behaviour: The Developmental Era tracks from the genesis of scholarly publication until 1945; The Initial Era from 1945 to 1980; and the Modern Era from 1980 to present day.

Figure 5 can be broken into three eras, with one phase aligned with the anecdotal epoch, and the second and third phases in the statistical epoch:

  • We define the “Developmental Era” as the period from 1665 to 1950, that appears to be split into two sections, separated by a sharp transition that occurs in 1798. As we will explain, the sharpness of this distinction is a result of the data in Dimensions (and of the systems on which Dimensions is based–specifically, systems that are based around the attribution of DOIs to academic articles [30]). Thus, we must regard the data reported until 1798 as unrepresentative of the actual development of scholarly publication and only anecdotal in nature.

    In the period until 1798 the Dimensions data contains just a few regularly publishing journals with total global number of articles being published reaching 288 in 1798. Up until that year, the journals with the largest publication volumes were Philosophical Transactions of the Royal Society of London, Transactions of the Linnean Society of London and Transactions of the Royal Society of Edinburgh, each of which were publishing around 30 articles per year and all of which had adopted a standardised title page, typically printed the names of authors in full form. But, in 1798, The Philosophical Magazine: A Journal of Theoretical, Experimental and Applied Physics was established by Alexander Tilloch of the London Philosophical Society, publishing 142 articles; the following year, in 1799, an even more voluminous journal launched as The British and foreign medical review was launched by the Royal College of Physicians, publishing 327 articles. Hence, there was a sharp and significant increase in the number of articles in Dimensions. These are the articles in publications that scholarly societies and other participants in the scholarly ecosystem decided were sufficiently valuable (and where means were sufficient) to clean up metadata and make DOIs or PubMed identifiers available. In the two significant journals mentioned here, author name forms are mixed in both journals with some articles being editorial (without author names), but while full form names still appear regularly, initial form becomes much more acceptable and this becomes established as a dominant signal entering the 1800s. Lefanu provides a for more detailed treatment of the rise of medical journals [29].

    In the years following 1800 the numbers of articles increase in volume to a level where some types of statistical analysis are appropriate, so long as the data are not filtered and segmented too finely. By 1823, the total, global number of journal articles tracked in Dimensions is consistently above 1000 (the level of a mid-sized research university today).

    It is worthy of note that throughout the whole of the Developmental Era there a low level of co-authorship and hence, we remain in Adams’ first age [19]. As we see in Fig. 2 it is not until 1958 that the modal average of authors on a paper moves from single-authorship to two co-authors.

    During this period scholarly communications are published in early versions of scholarly journals operated by early scholarly societies, with an intended audience of members, almost as we would treat newsletters today. Personal connections and members would often have been known to each other due to the small size of these early communities. Several innovations took place in this period, however, including the introduction of a somewhat standard title page to communications.

    In 1798, at the transition point, only 179 papers were published (as tracked by Dimensions under the constraint of Listing 1), but in 1799, 517 papers were published – a tripling of output. This significant increase in output appears to be coupled to the beginning of the proliferation of scientific journals. Although it is difficult to see the detail of this development in Fig. 6, which shows that there are currently around 70,000 actively publishing journals, in 1797 just four journals are recorded as active in publication in Dimensions (until just 1780, more than a century after the beginning of this movement, there had only been one or two active regularly publishing journals at any point in time, with 1781 being the first year when three journals published). From 1799, seven journals were regularly publishing. By 1827 this had become 19 and by 1837 it was 30. This part of significant growth began to see the development of norms and standards that needed to reach across journals and the community started to become large enough that the prior “gentlemanly” approach was no longer sufficient to handle the level of publication volumes and the community which needed to share in the knowledge being generated.

    During this period, the level of formality in the use of initial form versus full form in papers appears to have increased spontaneously and significantly.

  • The “Initial Era” began in the wake of the second world war, the dawn of Bush’s Endless Frontier [24] in 1945, and continues until 1980. It is the singular period in the history of the scholarly record in which the initial form dominated the full form of author names. The initial era is entirely contained in the statistical epoch and hence there are sufficiently many journals and papers that a statistical approach to analysis is valid.

  • The “Modern Era” has persisted from 1980 to the present day. It is an era in which we have seen the norm move steadily and rapidly toward the dominance of the use of full form names.

Figure 6: Development in the number of actively publishing journals.

Researcher-based analysis

In Fig. 5 we see the proportion of papers in which authors used initial form. In this section we ask a slightly different but related question, which is: How many authors used initial form? While one might expect this to be an identical question, and while it is related, there is a fundamental difference. Papers as objects exist in a world of journals, subject norms, publisher house styles and technological constraints. While each of these affects the authors, they have strong ties to national, cultural and linguistic effects. Thus, analyses of people rather than papers do add to our understanding of the developments shown above.

Figure 7 shows the globally conglomerated author view of the use of initial form on journal articles—it averages across national, cultural, gender, technological and disciplinary boundaries. We will decompose this plot in a variety of ways through the rest of the Results section, so it is important to gain familiarity with its overall shape and its key features. It is worthy of note that there is a an initial steep incline from 1945, marking the beginning of the Initial Era. The norm of initial-form dominance is maintained around 50% until around 1975 when the form starts being replaced by full form names. There is a rapid decrease in popularity of initial form between 1985 and 2000 and then, in around 2002 a precipitous drop, followed by a continued decline to around the 6% level is seen. It looks likely that the current decline will continue until initial-form is expunged entirely (or almost entirely).

Had we plotted Fig. 7 into the past, we would see that long-form names not only dominated the paper-based view but also the researcher-based view of publication in the anecdotal epoch.

Figure 7: Proportion of authors publishing with initial form names from 1945 to 2023.

Geographical analysis

The data in Fig. 7 are averaged across all countries but, over the time period that we are examining, not all countries participate in the research enterprise equally. Thus there is an implicit weighting in the global picture toward specific countries. If we speak about cultural, political, linguistic or geographic trends then it is important to understand the level of participation of different countries to an international average. As such, Fig. 8 reveals the contribution of each country to global paper production rates in Dimensions. The graph is produced by taking each paper and partitioning it by the countries of the affiliations of each author on the paper and then summing these contributions over all journal articles published in each year. In the plot, we have only plotted the top 13 countries and have conglomerated under the title ”Rest of World” as a 14th participant. As we can see from the Figure, during the period from 1945 to 1990, an extremely high percentage of global research output emanated from the US and the UK. While Japan and Germany make a significant and sustained contribution from the mid-1950s, between 50% and 60% of global output in the 1970s and 1980s came from the US and the UK – English-speaking countries with a somewhat shared cultural base. Thus, the behaviours summarised in Fig. 7 are dominated by an Anglo-American societal behaviours. However, as we see in Fig. 9 cultural alignment between the UK and the US may be more divergent than expected.

Figure 8: Percentage annual contribution to research journal publications by country of affiliation of author from 1945 to 2023.

Nonetheless, each country has its own social norms and traditions. Some countries are sufficiently large or diverse that there are different norms within the country while other countries have kinship with neighbours and there is some homogenisation between cultures. In our analysis here, we see some of these facets emerge naturally from the data.

The following three figures (Figs 910 and 11) show the proportion of authors publishing their name using initial form rather than full form—in each case around 25 countries are plotted (countries are selected to appear in the plot when they consistently participate in more than 500 papers per year) in grey with specific subsets picked out in colour.

Figure 9: Proportion of papers only displaying initial-form names broken down by countries of author affiliations from 1945 until present day. The grey dotted lines show all countries in the list for context, selected Western countries are picked out in colour. Fractional apportionment of papers to countries has been applied.

In Fig. 9 Australia, France, Germany, Spain, the United Kingdom and the United States are picked out in colours. We see that the United States is the country that most consistently leads on use of full form rather than initial form name forms on papers. Interestingly it shows a sharp jump in formality during the late 1980s followed by a slow reversion to its prior behaviour, ended by an almost symmetric sharp drop in the early 2000s. However, at all points, the US is the “familiar” country of the Western world. Australian authors have been the second most familiar since the late 1970s, having previously been slightly closer to their cousins in the United Kingdom, who having originally led on formality in this group in the 1940s through to the 1970s, started to be much more central in the group. While the Germans were highly informal in the 1940s, their increase in use of initial form rose significantly in the 1950s and 60s but has since settled close to the middle of this group. Finally, French authors have been probably more consistent than others in their embrace of the initial form, resulting in their moving from being one of the least formal to the most formal in the group while staying at more or less the same level of initial form usage between 1945 to 1990. They have then been consistently the highest user of initial form since the mid 1970s to present day. Of course, for all members of this group there has been a significant harmonisation over the last 20 years with extremely low levels of initial form usage in place today.

Figure 10: As Figure 9 with countries in the Asia-Pacific region and Brazil selected.

In Fig. 10 we have selected countries from the Asia-Pacific region and Brazil. In this case we see clearly the formalism (initial form alignment) of India from 1945 to present day with it being consistently (and in much of this time period significantly) above its comparators. Brazil passed the production threshold for this analysis in the late 1960s, at which point it was already more aligned to full name usage than most of the rest of the world (below most lines in the grey background), and continues to be even today. China, Japan and South Korea have all consistently been amongst the countries least likely to use initial form for the longest period of time. The facet in common for these countries is their use of non-Roman writing systems. Thus, in native language, author names are often single character, but when translated into English as choice is made to list the name as as a full name rather than just as a single initial. This may, in part, be due to cultural and linguistic factors in many Asian countries that give rise to common last names, the lack of a second or third given name, and hence the need to state full first names in order have the ability to disambiguate authors in a pre-ORCID environment [313233]. All these issues are compounded by transliteration in the scholarly record.

Figure 11: As Figures 9 and 10 with selected Slavonic-language countries picked out.

In Figs 9 and 10, it is notable that most countries that we reviewed have, at least in recent times, been below the modal behaviour regarding initial form usage. Thus, in Fig. 11 we explore the more formal part of this diagram, picking out Slavonic stages. Not all countries in this group speak a strictly Slavonic language, specifically, Romanian is an Eastern Romance language. Some countries in the group use the Roman alphabet while others use the Cyrillic alphabet, however, both of these alphabets are structured in a way so as to allow the delineation of first names and initials in a way that some of the Asian writing systems do not. However, the shared history of these countries (with the exception of Croatia) is that they formed part of the the orbit of the USSR before its fall in 1991.

A characteristic of these countries was the establishment of strong national academies, which may have led to greater harmonisation of norms of research culture at a national level. What is less clear is whether this is the core effect as this could also be due to pre-existing (i.e. non-research-specific) cultural and linguistic alignments. The data in Fig. 11 suggests that these countries have tended to be more formal in their use of initial form rather than full form in research discourse. Isolating the effect of national academies in research culture agenda-setting is beyond the scope of the current work. However, we conjecture that this may be a factor in determining behaviour. In the UK, the diffuse nature of the national academy (being split between the British Academy, the Royal Society and the Royal Academy of Engineering) has tended to mean that universities are more independent and powerful in their own right; in the US a similar diffuse system exists with a much more powerful university sector; and, in Germany the Max Planck Society, the Leibniz Association, Fraunhofer Society, Helmholtz Association and others provide a diverse setting for research culture to develop. Whether France’s more formal preference is due to linguistic structure or administrative structure is, again, hard to disinter—one may, after all, be a consequence of the other.

Figure 12: Proportion of papers published displaying only initial-form authors in 2020 versus power distance for a selection of countries. The size of each disk represents the cumulative volume of publications associated with the country between 1945 and 2023. Each disk is coloured according to the linguistic grouping of the principal language spoken in the country of the author’s affiliation.

Overall, over the last 20-year period we see a similar rate of harmonisation toward a more familiar norm. However, it is interesting to note a few of the journeys that we see in these data. For instance, it is noteworthy that Ukraine’s curve mirrors that of Russia throughout the period that we studied until circa 2014. Likewise, it is interesting to note that Belarus has continued to be consistently formal in its name format.

Europe is a particularly challenging setting for analysis of this nature due to the richness of the history of the region with significant changes in borders and definitions of countries over the period in which we are interested. In the Initial Era, some stability prevails but this is set against a backdrop of multinational history in which many countries in Europe have sub-populations of differing language and ethnicities – in some sense countries are an artificial construct. And yet, they are also the basis of national academies, evaluation systems, and systems of funding. Yet, at the same time, the Initial Era finds its genesis coincidentally at the time when the Marshall Plan was rebuilding Europe. A time of austerity for many countries in Europe, when there may have been a propensity toward more formal styles.

Beyond this, countries are also the places of education and the level of formality is instilled through institutions of learning and higher learning. Another confounding factor in our analysis is that in a more globalised world researchers are more mobile. Thus, the academic affiliation on their papers may not be indicative of their cultural, linguistic or educational background. Even more challenging, with the rise of international collaboration, it is unclear whether researchers will naturally import the style of other researchers with whom they have worked, who may not share their cultural background but who may convince them to adopt a different cultural norm.

Our data do seem to suggest some level of cultural, historical, linguistic or political coupling. For example, we see the close mirroring of practices between Ukraine and Russia until 2014, when political developments may have had cultural consequences, on the other hand we see a close relationship between Russia and Belarus seemingly not altering Belarus’s more conservative approach to publishing with a preference to initial form. Indeed, Belarusian-affiliated authors appear to be the sole example that do not follow the overwhelming international norm toward the use of full name on publications. However, it is very difficult to draw any solid conclusions that are not purely anecdotally-motivated from these data.

In an attempt to gain further insight, we explore the interplay of linguistic and cultural effects further by turning to the work of the cultural theorist Hofstede from the late 1960s. Through his work at IBM, Hofstede developed Cultural Dimension theory [343536] in an attempt to understand how cultural considerations affected organisational structures. One of the cultural dimensions that Hofstede developed, called the Power Distance Index, attempts to quantify the extent to which the less powerful members of organisations and institutions accept and expect power to be distributed unequally. As such, Power Distance may be thought of as a proxy for how formal a country is culturally in a professional setting. Of course, such things change with time and the Power Distance shown here is a snapshot at a particular point in time. We argue, however, that these things do not change as quickly from a cultural perspective as they do from a technological perspective.

Figure 12 plots the proportion of authors using an initial form aggregated at a country level against Hofstede’s Power Distance for that country. If there were to be a direct positive correlation between Power Distance and use of the initial form of the name on papers, then we would see country disks clustered close to a line that starts low and which gradually slopes upward to the right. However, in our plot we see little to no correlation. We see that most countries (including US, Canada, Japan, China et al.) in our selection are less formal (lower incidence of initial form) than the power distance in their country; while a few countries are just above the line (Australia, Germany, Hungary et al.), showing that they are slightly more formal than the power distance in their country suggests. Indeed, no country is significantly more formal in its usage of initial form then their power distance would indicate.

Countries with similar languages have similar power distances but there is little correlation with the proportion of authors using initial form. Indeed, countries with a first language of English have a Power Distance of around 0.4 but spread from 25% to 50% in their use of initial form; on the other hand Romance-language countries are associated with a wide range of power distances (50 for Italy to 90 for Romania) and yet there is a fairly consistent initial form percentage around 40%. However, Belarus, Russia and Ukraine are notably more formal (high power distance) and exhibit higher percentages of initial form usage (c. 70%) than other countries.

This suggests that formality of the cultural in which a piece of research is carried out is not a dominant factor in the adoption of initial form as a publication form. Indeed, since research is typically a highly collaborative endeavour we see people from different nationalities and cultures doing their work in different cultures; furthermore we see collaboration on single papers between different cultures and different locations. Nor does language or size of research output appear to have a particular effect. Nonetheless, the curves in Figs 910, and 11, share similar features and overall shapes in many cases – all the curves trend upward to a more formal style in the mid-1980s, all the curves fall back to less-formal levels after a peak around 1990 and many countries see a precipitous drop in formality in 2000. This similarity in shape suggests that there is a more fundamental underlying driver than a sociological one.

Disciplinary Analysis

Of course, sociology may differ by subject area more powerfully than by geography. When we grow up, we are taught certain cultural norms and practices, but the practices that we have for engaging in a research communication context are acquired at university either during undergraduate or postgraduate studies, depending on the field. In the post-1945 that we are investigating, university education had already become quite harmonised globally. With the world’s “first-mover” research nations had controlled and established a set of norms that new entrants have sought to emulate in order to engage on the same footing. (Figure 8 demonstrates the extent to which research would, de facto, take place in the English language and in the format chosen by the US, which in and of itself have been heavily influenced by its UK/European ancestry.) This included adoption of Bacon’s scientific method, the concept of the Humbolt institution and the format of scholarly communication in terms of the book and the journal article. In light of this harmonisation it is less surprising that we see correlations with national attitudes in publication practices and that it is more likely that we should see correlations elsewhere.

Figure 13: Evolution of contribution of different fields to the overall academic corpus with time. Attributions of papers to fields as per ANZSRC Field of Research Codes from Dimensions.

Figures 14 and 15 show analogous plots to Figs 910, and 11 but instead of breaking down the initial form percentage into country contributions, the basis has been changed to examine subject areas. In this case we make use of the 2020 ANZSRC Field of Research Codes that are automatically assigned to publications in Dimensions [37]. The plots must still, in aggregated form, reduce to Fig. 7 as was the case for the sum over all paths in the country-based plots and hence Fig. 13 shows the proportion of papers written in each of the subject areas and hence shows us the contribution level of each subject to the aggregate (or, the effective weighting of each line in each plot).

Figure 14: Proportion of journal articles in which authors only use initial form assigned to ANZSRC FoR-subject by year from 1945 to 2022. Selected Science, Technology, Engineering and Medicine (STEM) subjects are picked out in colour. Fractional apportionment is not applied – there is duplication of counting if papers cross 2-digit-FoR classifications. Fields are determined via Dimensionsautomated attribution to ANZSRC 2020 coding.

In the first subject-based Fig.  14, we have chosen to highlight STEM subjects. It is instantly noteworthy that these subjects cluster to the top of the graph, indicating a greater use of the initial form in the publications. The physical sciences appear to have the highest adoption of this format with information and computing sciences have the lowest level—both areas consistently hold these positions. Again, all fields have a similar overall shape, peaking in the late 1970s and early 1980s, since which the use of the initial form standard has declined. Interestingly, while the precipitous drop in 2000 still exists for some subjects the way that it is in the country data, a variety of fields do not show this feature, including Physical SciencesAgricultural, Veterinary and Food Sciences, and Engineering.

Figure 15: Proportion of journal articles in which authors only use initial form assigned to ANZSRC FoR-subject by year from 1945 to 2022. Selected Social Sciences, Humanities and the Arts for People and the Economy (SHAPE) subjects are picked out in colour. Fractional apportionment is not applied – there is duplication of counting if papers cross 2-digit-FoR classifications. Fields are determined via Dimensions automated attribution to ANZSRC 2020 FoR coding.

The second subject-based figure focuses on the SHAPE disciplines. With the exception of Built Environment and Design and, in a brief period form 1990-2000, for Philosophy and Religious Studies, all these fields exhibit lower levels of use of the initial form. As with the STEM-subject view, we see the 2020 precipitous change in the use of the initial form format is only significantly visible in Philosophy and Religious Studies, with Psychology showing this change to a lesser extent.

In C. P. Snow’s Tale of Two Cultures [38], he argues that Social Sciences, Arts and Humanities (which we will refer to as the SHAPE disciplines) have a different culture to Science, Technology, Engineering and Medicine (which we will refer to as the STEM disciplines), and this appears to be born out in our data.

With the exception of Built Environment and, Philosophy and Religious Studies for a brief period in the 1990s, the SHAPE disciplines lie below the STEM disciplines consistently in their use of initial form. When we take account of the weighting of volumes in different disciplines shown in Fig. 13 then it is clear that the average of disciplines will be weighted toward the disciplines of Biomedical and Clinical Sciences and Health Sciences and Biological Sciences. All three areas are aligned with medicine and hold a relationship to the formality associated with healthcare professionalism. At the same time they are three of the largest disciplines by number of outputs, accounting for 40%-50% of global publications.

Journal analysis

It is of course impossible to consider the evolution of articles, without also considering the evolution of practices at the journal level. Journals have historically been associated with scholarly societies that are disciplinary in nature and have embodied communities and have been reflections of their practice. Not only this, journals are the wrapper for article publications with journal management teams having responsibility for technological choices such as how the move from print to online was accomplished and when it took place. In the digital era, they have also held the choice of specific platforms and hence have had their options defined for them by suppliers of these technologies. Not limited to the digital era of publishing, journal editors have made choices such as the adoption of policies such as house styles which opine on the use of grammar, and of author name styling. Understanding these effects in situ is a key input to further enhance our picture of the landscape.

Figure 16: Proportion of initial form usage in three journals, Journal of Biological Chemistry, Tetrahedron and The BMJ from 1950 to 2022.

Figure 16 shows the changes that we see in Fig. 12 at the micro level. The examples chosen in Fig. 16 are well documented and hence we can see some of the changes that took place through the Initial Era with more of the detail.

In the mid nineties, Highwire press emerged as a significant player in the online hosting of Journal content. Starting with the Journal of Biological Chemistry, Highwire was at the forefront of re-imagining the definition of a paper from a digital ‘photocopy’ of a physical paper, Highwire’s technology allowed the paper to be a fully interactive digital object in a contemporary sense—text was textual rather than graphical and hence was searchable, elements of the body of the document were broken down into component parts so that images and their captions were distinct objects in the paper [3940].

By 2000, the Journal of Biological Chemistry noted that, following the lead of Highwire press, many journals now considered the online version of the publication to be the version of record.[2] As Fig. 16 shows, the precipitous switch from initial-form to full-form names in 1995 coincides precisely with the launch of the journal online, suggests that the new system made first names mandatory.

The BMJ makes a similar shift, also in 1995, although its profile suggests that that first names are strongly preferred but not enforced [41]. Prior to this, in around 1975, it appears that all three journals, updated their editorial policies as there are sharp changes of behaviour. Both the BMJ and the Journal of Biological Chemistry appear to have put journal policies in place that mandate (or close to mandate) initial form, while Tetrahedron appears to have either removed a mandate for initial form or even actively promoted full name form given the initial rapid drop off from close to 100% to 60% initial-form name usage within just five years from 1975-1980.

By 2016, the BMJ had switched to using ScholarOne for manuscript submission. As demonstrated in the BMJ tutorial, within this workflow author names are not keyed in individually, but like a CRM system, selected from a database of previously recorded identities. If a person cannot be found (by email address), a new record can be created, and given names and last names are required. The system is normative in the sense that no indication is given that a given name could just be an initial. Authors are not so much recorded against the article as people are linked. A single representation of an author is now used across all journals that use the ScholarOne service. The BMJ sees no usage of initial-form names from 2016.

The effect of technology as a driver of change is clearly delineated in this plot as Tetrahedron declines gradually indicating a change in social preference or weaker forms of change such as journal policy and house style changes. Tetrahedron did not go live with its first online submission system until 1995, so it is assumed that this change was managed offline by typesetters.

Technological Analysis

We may reasonably ask what the impetus was for the change in behaviour in the medicine-related fields seen in a previous section (the significant, sharp discontinuity that we see for Biomedical and Clinical Sciences, and Health Sciences in Fig. 14). The speed of change suggests that this was a technological rather than a cultural driver, since technological changes tend to be implemented with greater speed. Indeed, since the medical fields share infrastructures such as MedLine and PubMed, which are funder orientated, there could be a strong funding-aligned impetus that we can pinpoint in this case.

In order to explore PubMed we need to find differentiating characteristics of the PubMed data to be able to track effects. Helpfully, PubMed and its anticedents such as MedLine have conferred unique identifiers to articles prior to the advent the generalised use of Crossref DOIs. This means that some articles on PubMed do not have DOIs and only have a PubMed identifier.

Figure 17: Percentage incidence of authors using initial form on PubMed articles with (blue) and without DOIs (red).

Figure 17 shows distinct and different behaviours between authors associated with articles appearing on PubMed that are associated with DOIs and those which aren’t. The blue line shows the percentage of authors associated with a DOI where initial form is used; the red line shows the same percentage for authors associated with articles without a DOI. For articles with a DOI, we see a gradual change to lower rates of initial form being used with authors migrating to full form. We note a small drop in the blue line around 2002 which, we speculate, would be associated with the launch of PubMed’s new platform. This new platform supported given name metadata fields and hence naturally encouraged the use of full form names [4243].

Articles without a DOI (only having a PubMed identifier), shown in the red line of Fig. 17, have a distinctly different shape. The difference in behaviour of the red line is explained by a systemic development: When Crossref introduced the opportunity to implement DOIs for academic articles, their metadata format was sufficiently fully formed and modern to allow for a field to include given name and publishers adopting the new standard were able to update their back catalogues with the additional data. Thus, the blue line is a representation of the actual metadata included on articles as they were published, whereas the red line is the result of the legacy metadata standard that existed prior to the advent of Crossref’s metadata schema [44]. Even though articles without a DOI may, in fact, have full form names on the publication this is not represented in the metadata.

Similar parallels in behaviour can be seen with other major bibliographic and bibliometric databases that existed through this period, with Web of Science making collecting first names from records processed after 2007, and allowing them to be searched in 2011 [45], which Scopus had done in 2004 [4647].

The ongoing decline in the use of initial form from 2002 to present day, may not be due only to technology changes but rather may be motivated by changes to another scholarly institution—research evaluation. In the period from 2002, technological developments took place to introduce of (current research information systems) CRISes [48], to address the need for reporting data in research economies that introduced research evaluation such as the UK and Australia. It was not until 2010 the ORCID launched formally to tackle the disambiguation problem at a deeper systemic level – perhaps surprisingly the launch of ORCID does not seem to have a particularly direct effect on these data [495051].

Gender Analysis

Finally, we turn our attention towards a gender-focused analyses. We have pointed out reasons that gender may be hidden by choice in certain contexts in the introduction to this paper. However, it is clear that the reasons for the obscuring of gender in research publication are considerably more subtle [5253]. Countless studies demonstrate differences in outcome for women in peer review in both grant and publication contexts, and citations of their work is lower than for male counterparts [54555628]. Further work seeks to understand the causal nature of these effects [57]. What we show here is that our data analysis as a further dimension for consideration namely that, perhaps unsurprisingly, technology can be a significant modifier to gender visibility in publication. We note that the results shown in this section rely on statistical methods and hence have known limitations, however, we do believe them to be sufficiently robust for the purposes of a high-level indicative analysis, as presented here. The work of Lockhart et al.[58] provides a critical understanding of some of the technical issues with this type of analysis.

Figure 18: The development of perceived gender of authors on PubMed publications from 1980 to 2022. Blue line: Authors with names statistically associated with men; Red line: Authors with names statistically associated with women; Grey line: Authors with initials only; Green line: Authors with statistically indeterminate gendered-names.

Figure 18 illustrates the development of the perception of gender in authorship from 1980 to 2022 in the context of PubMed. As we have seen in Fig. 17 there is a distinct technological event in 2002 that caused a shift away from initial form and toward full form names. We see that both the blue line (names statistically associated with men) and the red line (names statistically associated with women) jump upward around 2002, at which point there is a steep decline in initial form usage (grey line). Increased representation of full names has given rise to an increase in indeterminacy (green line), which has steadied at around 10% of output.

While the step in the blue line in 2002 draws the eye it is important to take this change as a proportion of the population—a move from circa 38% names statistically associated with men to circa 44% is a 15.7% increase. However, for names statistically associated with women, the move is from 12% to more than 14% – closer to a 20% increase. In addition the gradient of the increase in names statistically associated with men in the 20 years from 2002 to 2022 is at most 8% (from 44% to 52%) compared with names statistically associated with women which has moved by more than 10%.

While the reporting of gender participation in research is still relatively nascent with systematised aggregation of reporting only available since around 1996. Already in many regions of the world that reported in 2002, the date of the change in PubMed systems, women regularly accounted for between 30% and 50% of researchers [59]. Yet, PubMed authorships at the time attribute only 15% of output to those with names statistically associated with women, while around 32% of output was either of indeterminate or unknown gender. This, together with the steeper rise in participation of women in using full form names, suggests that the initial form convention proportionally over represented women over men, making the contributions of women researchers systematically more difficult to identify.

Discussion

We have argued that the period from 1945 until 1980 defined a period that we call the “Initial Era”. Rather than being a normal period, we believe that it is exceptional in the scholarly record, being the only era that we can identify in which there is an extended period where initial form was used instead of full form names on academic manuscripts.

Each of the subsections in the Results section of this paper examines the data from a different aspect to allow us to, firstly, assess the robustness of the data, and secondly, to draw together a picture of the causes and effects of the Initial Era. In this section we will use our results to speculate on the likely genesis of the Initial Era and to examine the effects. We will finish by considering the future of the scholarly record and the potential for future bibliometric archaeologists to perform similar studies.

The rise of the Initial Era

We speculate that the adoption of initial form during this exceptional 40-year period is far from Vannevar Bush’s Endless Frontier [24] being accidental, as it was a highly influential document that appeared at the dawn of US-preeminence, a time when science was being established as “serious business” with a more formal cultural norm associated with it. This may have been influenced by the rise in importance of science and medicine and the view that technology could solve all problems. Similarly, the Marshall Plan rebuilding of Europe could have led to the import of norms the US, but this seems less likely since the US is continually one of the lower users of initial-form (see Fig. 9).

However, during this period, research output was dominated by the US, the UK and Japan—all countries that were predisposed to publish in English and to reflect the cultural norms in the US for political reasons. The Bretton-Woods Era came to an end in the late 1970s and the Cold War in the late 1980s, and with the rise of globalisation greater international collaboration has been possible, leading to Adam’s Fourth Age [19]. Yet, as we see from Figs. 910 and 11, the use of initial form persisted (and even increased, possibly due to the constraints of paper media and their interaction with the era of Big Science) from 1980 to 2000. As demonstrated in our technological analysis, it was a combination of house style decisions, technology changes in the form of PubMed and Crossref, and the switch to online journal submission systems that fundamentally changed research culture back to full name usage. It remains a “chicken and egg” issue as to whether technology evolved to meet cultural needs or whether technological change gave cultural norms an opportunity to assert themselves.

The methodology behind our analysis in Sec. III.3 is prone to reflecting international trends. By construction, the unit of analysis is the paper and is classified as an initial-form paper only in the situation that all authors use that form. There are normative trends implicit in a global research community—publishers have house styles that are applied regardless of origin; if the authors who brought funding for the work, or who are in some other way senior, choose a style of publishing their name on the paper perhaps there is a psychological bias that goes on that we cannot track in the data; if a majority of authors choose a specific style then how many authors are comfortable breaking step and choosing a different form for their name? International collaboration is another anchor to which our methodology is particularly sensitive as only papers that are entirely initial-based factor in analysis such as Sec. III.3—hence, a core assumption of the analysis is that the kinds of effects discussed here were sufficiently strong as to normalise behaviour across a paper regardless of geographic background.

As collaboration has become more of a norm in research (see Fig. 3) and as many countries have invested in and developed their own advanced research economies (see Fig. 13), a much more diverse community has emerged. Hence, once “senior partners” are now just simple partners. On the one hand, this will necessarily require the representation of newly evolved norms. Yet, the norms are already set and new norms take time to evolve.

The fall of the Initial Era

As with any era in history, a critical understanding the outcomes and impacts of the period are critical in understanding ourselves and improving current and future practices.

The fall of the Initial Era is fairly clear cut from a technological perspective. We posit that the current trend is a mix three major factors:

  1. 1. Technological harmonisation: Allowing the choice of capturing and using full names across many different types of system from research information management within institutions to evaluation, and from publishing to grant application and administration;

  2. 2. Globalisation of research: The natural harmonisation of norms as more and more researchers collaborate outside a narrow context;

  3. 3. Trust: The ongoing need to propagate trust in the sense of understanding the identity of participants in the research ecosystem.

All three of these points play into the greater narrative of how research is changing around us. There is a greater need for transparency as research has an increasingly significant effect on everyday lives and as governments investment more in research, it is critical to make knowledge common. The move to greater transparency is multifaceted—making the knowledge itself more transparently available is just one stream; making its production and provenance transparent is another. Without transparency of production and provenance, we cannot expect the trust, which is needed for research to progress.

The articulation of this need for openness is often found in evaluation mechanisms, either national or funder based. All the factors that we identify above relate to evaluation. Technological developments, especially around name disambiguation, relate to the rise of ORCID as a standard [6061], the use research information management systems and research profiling tools [62]. The globalisation of research has led to increased researcher movement and for researchers to be able to move freely, they need to be able to assert their credentials—which involves asserting identity in relation to their research work. And, in a world where nefarious actors are on the rise [63], understanding the provenance of research often relies on understanding identity.

While one may think that the study of the use of names in academic work is perhaps facile, our study shows that it is anything but trivial. Ultimately, we believe that it is possible to justify the argument that the effect of a period in which the usage of given names is suppressed, during the Initial Era, has been to hide and marginalise the contribution of women to research; suppress data about social norms and increase the vulnerability of the research system to abuse by weakening provenance.

The future of bibliometric archaeology

If we have learned anything in the writing of this paper, we must reflect that the lens through which we look is as important as the analysis itself. While Dimensions is a useful tool in understanding demographics and evolution of certain types of literature, it is a product of its world and of its construction [30]–specifically the requirement of a DOI or other mainstream identifier limits Dimensions’ coverage at this time. Data coverage pre-1800 shows low volumes of output and, in part, this is because output levels are lower but it is also because strategic choices have had to be made by infrastructure organisations in what should be afforded, the cost and effort of creating digital versions and digital identifiers for past material that is not likely to be of current research value in a non-sociohistorical setting. This mirrors the more hidden lack of data that we have identified in the Initial Era, where the contribution of women is hidden through the technical implementation of a social construction.

In the current contemporary setting, capturing and retaining information that describes an individual is changing: GDPR regulations in Europe and their equivalents around the world are designed to allow people to be forgotten if they wish. Yet understanding the identity of those who carried out a piece of research seems natural to many of us and even essential information in order to understand biases that may be present. If one is thinking purely scientifically about record keeping, then it makes sense to gather as much data as possible. However, this shows a profound lack of subtlety and even respect when considering issues of identity and ways of knowing—cultural approaches that come from a different tradition than that of Western research but which, as research diversifies its community, need to be included.

It is apposite to ask whether we live in a special time in bibliometric analysis where name data continues to be freely available and whether this period will persist. ORCID now assigns unique identifiers to any researcher who would like to have one. We live in a time of increasing polarisation and researchers who wish to carry out certain types of research or express particular opinions can face censure or removal of funding. If we take the logic of using research identifiers further, some might argue that the adoption of cryptographic identities, like Satoshi Nakamoto, would allow for researchers to regain intellectual freedom. Zero-knowledge proofs could be used to authenticate claims for career progression, double-blind peer review would be built in to such a system, potentially increasing the overall fairness of both publication and funding practices. Others might argue that this undermines deeply the trust that we have in the scholarly record and permits abusive expression or propagation of fake research or support of political or commercial interests from “behind a mask”. This is, perhaps, the modern-day equivalent of the binaries pointed out by Evans et al. [64] in their commentary on the role of the algorithm.

From a purely analytical standpoint, it would be sad to see the scholarly record lose an aspect that is reflective of current trends. But, perhaps one role for blockchain-like technologies may be in the detailed capture of demographic information of researchers in a way that protects it for future generations in much the same way that detailed census results are not revealed at the time but which are held in trust for a hundred years. While statistical information on gender, culture and background might be helpful in workforce planning and in helping to ensure that the research system carries on working toward greater levels of inclusion and representation, detailed information may be preserved for future historians and bibliometric archaeologists.

More positively, it could be argued that the transformation away from name formalism should not stop at author bylines. Name formalism is also embraced in reference formats. It could be argued that even within a paper, this formalism suppresses the diversity signal in the research that we encounter. Reference styles were defined in a different era with physical space constraints. Is it time to reconsider these conventions – establishing new paths of analysis in the bibliographic record?

Within contribution statements that use the Credit Ontology [65], initials are also commonly employed to refer to authors although this is not part of the standard. This convention also creates disambiguation issues when two authors share the same surname and first initials. Here too, as the Digital Structure of a paper continues to evolve, we should be careful not to unquestioningly embed the naming conventions of a different era into our evolving metadata standards.

Acknowledgements

The authors wish to thank Briony Fane for her careful reading and commenting on the manuscript, and to Kathryn Weber Boer for her suggestions to improve discussions on issues of gender.

Data Availability

The data and code for this paper is available from Figshare at: https://doi.org/10.6084/m9.figshare.25664154

Author Contribution

Simon J Porter: Conceptualization; Formal Analysis; Methodology; Visualisation; Writing – original draft; Writing – review & editing.

Daniel W Hook: Formal Analysis; Methodology; Visualisation; Writing – original draft; Writing – review & editing.

Conflict of interests

The authors of this paper are both employees of Digital Science, the owner and operator of Dimensions.

References

  1. P. A. Hall, ed., Varieties Of Capitalism: The Institutional Foundations of Comparative Advantage, illustrated edition ed. (Oxford University Press, U.S.A., Oxford England ; New York, 2001).

  2. J. Jarvis, The Gutenberg Parenthesis: The Age of Print and Its Lessons for the Age of the Internet (Bloomsbury Academic, New York, 2023).

  3. A. Johns, The Science of Reading: Information, Media, and Mind in Modern America (University of Chicago Press, Chicago; London, 2023).

  4. M. Taylor, “How do publications find their audience?” (2024), publisher: figshare.

  5. B. Macaluso, V. Larivière, T. Sugimoto, and C. R. Sugimoto, Academic Medicine 91, 1136 (2016).

  6. C. Ni, E. Smith, H. Yuan, V. Larivière, and C. R. Sugimoto, Science Advances 7, eabe4639 (2021), publisher: American Association for the Advancement of Science.

  7. C. R. Sugimoto, Y.-Y. Ahn, E. Smith, B. Macaluso, and V. Larivière, The Lancet 393, 550 (2019), publisher: Elsevier.

  8. D. Kozlowski, D. S. Murray, A. Bell, W. Hulsey, V. Larivière, T. Monroe-White, and C. R. Sugimoto, PLOS ONE 17, e0264270 (2022a), publisher: Public Library of Science.

  9. D. Kozlowski, V. Larivière, C. R. Sugimoto, and T. Monroe-White, Proceedings of the National Academy of Sciences 119, e2113067119 (2022b), publisher: Proceedings of the National Academy of Sciences.

  10. A. Meadows, Journal of Information Science 11, 27 (1985).

  11. D. W. Hook and S. J. Porter, Frontiers in Research Metrics and Analytics 6 (2021).

  12. L. Bornmann, R. Haunschild, and R. Mutz, Humanities and Social Sciences Communications 8, 1 (2021), number: 1 Publisher: Palgrave.

  13. Philosophical Transactions of the Royal Society of London 4, 1075 (1997), publisher: Royal Society.

  14. L. A. Whited, ed., The Ivory Tower and Harry Potter: Perspectives on a Literary Phenomenon (University of Missouri Press, Columbia, Mo., 2004).

  15. S. Nakamoto, “Bitcoin: A Peer-to-Peer Electronic Cash System,” (2008).

  16. R. E. Sorge, L. J. Martin, K. A. Isbester, S. G. Sotocinal, S. Rosen, A. H. Tuttle, J. S. Wieskopf, E. L. Acland, A. Dokova, B. Kadoura, P. Leger, J. C. S. Mapplebeck, M. McPhail, A. Delaney, G. Wigerblad, A. P. Schumann, T. Quinn, J. Frasnelli, C. I. Svensson, W. F. Sternberg, and J. S. Mogil, Nature Methods 11, 629 (2014), publisher: Nature Publishing Group.

  17. A. Katsnelson, Nature (2014), 10.1038/nature.2014.15106, publisher: Nature Publishing Group.

  18. S. Reardon, Nature (2017), 10.1038/nature.2017.23022, publisher: Nature Publishing Group.

  19. J. Adams, Nature 497, 557 (2013), number: 7451 Publisher: Nature Publishing Group.

  20. M. Thelwall and N. Maflahi, Quantitative Science Studies 3, 331 (2022).

  21. The first paper in Dimensions with co-authorship of more than 100 authors is in Chemistry from 1928 |66.

  22. J. P. Pardo-Guerra, The Quantified Scholar: How Research Evaluations Transformed the British Social Sciences (Columbia University Press, 2022).

  23. S. J. Porter and D. W. Hook, Frontiers in Research Metrics and Analytics 7 (2022).

  24. V. Bush, The Endless Frontier, Report to the President on a Program for Postwar Scientific Research, Tech. Rep. (OFFICE OF SCIENTIFIC RESEARCH AND DEVELOPMENT WASHINGTON DC, 1945).

  25. J. R. Oppenheimer, Proceedings of the National Academy of Sciences 50, 1194 (1963), publisher: Proceedings of the National Academy of Sciences.

  26. D. Yan, Z. Hai-qing, D. Guo-bao, and T. Cheng, Chinese Journal of Integrative Medicine 11, 229 (2005).

  27. O. Brück, Communications Medicine 3, 1 (2023), publisher: Nature Publishing Group.

  28. A. Holmes and S. Hardy, “Gender bias in peer review — opening up the black box,” (2019).

  29. W. R. Lefanu, Bulletin of the Institute of the History of Medicine 5, 735 (1937), publisher: The Johns Hopkins University Press.

  30. D. W. Hook, S. J. Porter, and C. Herzog, Frontiers in Research Metrics and Analytics 3 (2018).

  31. W. A. Kretzschmar, in Language in the USA: Themes for the Twenty-first Century, edited by E. Finegan and J. R. Rickford (Cambridge University Press, Cambridge, 2004) pp. 39-57.

  32. T. A. Velden, A.-u. Haque, and C. Lagoze, in Proceedings of the 11th annual international ACM/IEEE joint conference on Digital libraries, JCDL ’11 (Association for Computing Machinery, New York, NY, USA, 2011) pp. 241-250.

  33. Y. Liu, L. Chen, Y. Yuan, and J. Chen, American Journal of Physical Anthropology 148, 341 (2012).

  34. G. Hofstede, Journal of International Business Studies 14, 75 (1983).

  35. G. Hofstede, Culture’s Consequences: Comparing Values, Behaviors, Institutions and Organizations Across Nations, second edition ed. (SAGE Publications, Inc, Thousand Oaks, Calif., 2003).

  36. G. Hofstede, Online Readings in Psychology and Culture 2 (2011), 10.9707/2307-0919.1014.

  37. S. J. Porter, L. Hawizy, and D. W. Hook, Quantitative Science Studies 4, 127 (2023).

  38. C. P. Snow, The Two Cultures, reissue edition ed. (Cambridge University Press, Cambridge, 2012).

  39. Journal of Biological Chemistry 275, 12361 (2000).

  40. A. B. Watson, Journal of Vision 11, ii (2011).

  41. J. Smith and R. Smith, BMJ 312, 1626 (1996).

  42. NLM, “MEDLINE Data Changes – 2002. NLM Technical Bulletin. Nov-Dec 2001,” (2002).

  43. NLM, “Skill Kit: Searching Full Author Names in PubMed®,” (2009), publisher: U.S. National Library of Medicine.

  44. Crossref, The Formation of CrossRef: A Short History, Tech. Rep. (CrossRef, 2009).

  45. Clarivate, “Web of Science Core Collection: Explanation on Full Author Names,” (2022).

  46. P. J. Hane, “Elsevier Announces Scopus Service,” (2004).

  47. E. Sullo, Journal of the Medical Library Association 95, 367 (2007).

  48. B. Jörg, Data Science Journal 9 (2010), 10.2481/dsj.CRIS4.

  49. D. Butler, Nature 485, 564 (2012), number: 7400 Publisher: Nature Publishing Group.

  50. Q. Schiermeier, Nature 526, 281 (2015), number: 7572 Publisher: Nature Publishing Group.

  51. A. Meadows, L. L. Haak, and J. Brown, 32, 9 (2019), number: 1 Publisher: UKSG in association with Ubiquity Press.

  52. C. López Lloreda, (2022), 10.1126/science.caredit.adf3063.

  53. E. G. Teich, J. Z. Kim, C. W. Lynn, S. C. Simon, A. A. Klishin, K. P. Szymula, P. Srivastava, L. C. Bassett, P. Zurn, J. D. Dworkin, and D. S. Bassett, Nature Physics 18, 1161 (2022), publisher: Nature Publishing Group.

  54. C. D. Zhou, M. G. Head, D. C. Marshall, B. J. Gilbert, M. A. El-Harasis, R. Raine, H. O’Connor, R. Atun, and M. Maruthappu, BM) Open 8, e018625 (2018), publisher: British Medical Journal Publishing Group Section: Health policy.

  55. H. Draux and S. Kundu, Gender Imbalance in Cancer Research Grants, Tech. Rep. ([object Object], 2018) artwork Size: 2013631 Bytes.

  56. H. Draux, S. Kundu, and S. Porter, Gender Representation in UK Research Institutions, Tech. Rep. ([object Object], 2019) artwork Size: 3057372 Bytes.

  57. V. A. Traag and L. Waltman, (2022), arXiv:2207.13665 [cs.DL]

  58. J. W. Lockhart, M. M. King, and C. Munsch, Nature Human Behaviour 7, 1084 (2023).

  59. U. I. for Statistics, “Share of women among total researchers by country, 1996-2018,”

  60. L. L. Haak, A. Meadows, and J. Brown, Frontiers in Research Metrics and Analytics 3 (2018).

  61. S. J. Porter, Frontiers in Research Metrics and Analytics 7, 779097 (2022).

  62. V. Ilik, M. Conlon, G. Triggs, M. White, M. Javed, M. Brush, K. Gutzman, S. Essaid, P. Friedman, S. Porter, M. Szomszor, M. A. Haendel, D. Eichmann, and K. L. Holmes, Frontiers in Research Metrics and Analytics 2 (2018), 10.3389/frma.2017.00012, publisher: Frontiers.

  63. S. J. Porter and L. D. McIntosh, “Identifying Fabricated Networks within Authorship-for-Sale Enterprises,” (2024), arXiv:2401.04022 [cs].

  64. J. Evans, T. Reigeluth, and A. Johns, Osiris 38, 19 (2023).

  65. L. Allen, J. Scott, A. Brand, M. Hlava, and M. Altman, Nature 508, 312 (2014).

  66. F. Emich, A. Benedetti-Pichler, F. Henrich, L. Moser, R. Strebinger, F. Pregl, F. Zaribnicky, L. Rosenthaler, E. M. Chamot, H. Herbst, H. Fitting, H. Ambronn, A. Frey, L. Wright, K. John, F. K. Reinsch, A. Zimmern, M. Coutin, E. Lehmann, J. L. Pech, R. Harder, H. Siedentopf, C. Spierer, Lieberkühn, W. Kaiser, Rheinberger, W. Kraemer, C. Kern, H. Pohle, S. v. Wachenfeldt, F. Lossen, A. Köhler, Proell, O. Linde, E. Saxl, W. Schäffer, J. Kisser, F. K. Studnióka, F. Roll, T. Huzella, R. Chambers, T. Péterfi, C. V. Taylor, G. de Mottoni, D. L. Parkhurst, H. Utermöhl, Goring, E. Naumann, Volk, G. Lunde, E. A. Hill, E. Q. Adams, G. Linzenmeier, E. Kaufmann, P. Kirkpatrick, M. C. Magarian, R. Wolff, K. Schuhecker, H. C. Hagedorn, B. N. Jensen, W. Geilmann, R. Höltie, L. Dienes, L. Pincussen, P. B. Rehberg, S. Wermuth, M. Shepherd, H. B. Rasmussen, C. E. Christensen, R. Mellet, M. A. Bischoff, H. Hiller, J. J. Hopfield, F. Paneth, K. Peters, P. Günther, H. M. Elsey, M. Crespi, E. Moles, W. Kliefoth, O. E. Frivold, R. E. Burk, B. Noyes, W. Ewald, R. Whytlaw-Gray, H. Whitaker, H. Figour, F. Sautier, F. Verzár, J. Barcroft, L. Condorelli, E. P. Poulton, W. R. Spurrell, E. C. Warner, R. Suhrmann, K. Clusius, J. J. Manley, W. A. Roth, G. Naeser, O. Döpke, H. Leontjew, H. S. Patterson, R. W. Gray, R. A. Millikan, H. G. Barbour, W. F. Hamilton, Dickinson, E. A. Vuilleumier, W. Klemm, W. Biltz, A. Stock, and G. Ritter, Zeitschrift für analytische Chemie 74, 191 (1928).

Editors

Ludo Waltman
Editor-in-Chief

Ludo Waltman
Handling Editor

Editorial assessment

by Ludo Waltman

DOI: 10.70744/MetaROR.15.1.ea

This article presents a large-scale data-driven analysis of the use of initials versus full first names in the author lists of scientific publications, focusing on changes over time in the use of initials. The article has been reviewed by three reviewers. The originality of the research and the large-scale data analysis are considered strengths of the article. A weakness is the clarity, readability, and focus of certain parts of the article, in particular the introduction and background sections. In addition, the reviewers point out that the discussion section can be improved and deepened. The reviewers also suggest opportunities for strengthening or extending the article. This includes adding case studies, extending the comparative analysis, and providing more in-depth analyses of changes over time in policies, technologies, and data sources. Finally, while reviewer 3 is critical about the gender analysis, reviewer 2 considers this analysis to be a strength of the article.

Competing interests: None.

Peer review 1

Dmitry Kochetkov

DOI: 10.70744/MetaROR.15.1.rv1

The presented preprint is a well-researched study on a relevant topic that could be of interest to a broad audience. The study’s strengths include a well-structured and clearly presented methodology. The code and data used in the research are openly available on Figshare, in line with best practices for transparency. Furthermore, the findings are presented in a clear and organized manner, with visualization that aid understanding.

At the same time, I would like to draw your attention to a few points that could potentially improve the work.

  1. I think it would be beneficial to expand the annotation to approximately 250 words.
  2. The introduction starts with a very broad context, but the connection between this context and the object of the research is not immediately clear. There are few references in this section, making it difficult to determine whether the authors are citing others or their own findings.
  3. The transition to the main topic of the study is not well-defined, and there is no description of the gap in the literature regarding the object of study. Additionally, “bibliometric archaeology” appears at the end of the introduction but is only mentioned again later in the discussion, which may cause confusion for the reader.
  4. It would be helpful to clearly state the purpose and objectives of the study both in the Introduction and in the abstract as well.
  5. Besides, it is important to elaborate on the contribution of this study in the introduction section.
  6. The same applies to the background – a very broad context, but the connection with the object of the research is not entirely clear.
  7. Page 4 – as far as I understand, these are conclusions from a literature review, while point 3 (Reflective Richness of Data) does not follow from the previous analysis.
  8. The overall impression of the introduction and background is that it is an interesting text, but it is not well related to the objectives of the study. I would recommend shortening these sections by making the introduction and literature review more pragmatic and structured. At the same time, this text could be published as a standalone contribution.
  9. As I mentioned above, the methodology refers to the strengths of the study. However, in this section, it would be helpful to introduce and justify the structure of presenting the results.
  10. In the methodology section, the authors could also provide a footnote with a link to the code and dataset (currently, it is only given at the end).
  11. With regard to the discussion, I would like to encourage the authors to place their results more clearly in the academic context. Ideally, references from the introduction and/or literature review would reappear in this section to help clarify the research contribution.
  12. Although Discussion C is an interesting read, it seems more related to the introduction than the results. Again, the text itself is rather interesting, but it would benefit from a more thorough justification.

Remarks on the images:

  1. At least the data source for the images should be specified in the background, because it is not obvious to the reader before describing the methodology.
  2. The color distinction between China and Russia in Figure 8 is not very clear.
  3. The gray lines in Figures 9-11 make the figures difficult to read. Additionally, the meaning of these lines is not clearly indicated in the legends of Figures 10 and 11. These issues should be addressed.

All comments and suggestions are intended to improve the article. Overall, I have a very positive impression of the work.

Sincere,

Dmitry Kochetkov

Competing interests: None.

Peer review 2

Erjia Yan

DOI: 10.70744/MetaROR.15.1.rv2

Overview

This manuscript provides an in-depth examination of the use of initials versus full names in academic publications over time, identifying what the authors term the “Initial Era” (1945-1980) as a period during which initials were predominantly used. The authors contextualize this within broader technological, cultural, and societal changes, leveraging a large dataset from the Dimensions database. This study contributes to the understanding of how bibliographic metadata reflects shifts in research culture.

Strengths

+ Novel concept and historical depth

The paper introduces a unique angle on the evolution of scholarly communication by focusing on the use of initials in author names. The concept of the “Initial Era” is original and well- defined, adding a historical dimension to the study of metadata that is often overlooked. The manuscript provides a compelling narrative that connects technological changes with shifts in academic culture.

+ Comprehensive dataset

The use of the Dimensions database, which includes over 144 million publications, lends significant weight to the findings. The authors effectively utilize this resource to provide both anecdotal and statistical analyses, giving the paper a broad scope. The differentiation between the anecdotal and statistical epochs helps clarify the limitations of the dataset and strengthens the authors’ conclusions.

+ Cross-disciplinary relevance

The study’s insights into the sociology of research, particularly the implications of name usage for gender and cultural representation, are highly relevant across multiple disciplines. The paper touches on issues of diversity, bias, and the visibility of researchers from different backgrounds, making it an important contribution to ongoing discussions about equity in academia.

+ Technological impact

The authors successfully connect the decline of the “Initial Era” to the rise of digital publishing technologies, such as Crossref, PubMed, and ORCID. This link between technological infrastructure and shifts in scholarly norms is a critical insight, showing how the adoption of new tools has real-world implications for academic practices.

Weaknesses

– Lack of clarity and readability

While the manuscript is rich in data and analysis, it can be dense and challenging to follow for readers not familiar with the technical details of bibliometric studies. The text occasionally delves into highly specific discussions that may be difficult for a broader audience to grasp while other concepts are introduced in cursory. Consider condensing the introduction section, removing unrelated historical accounts, and leading the audience to the key objectives of this research much earlier.

– Missing empirical case studies

The manuscript remains largely theoretical, relying heavily on data analysis without providing concrete case studies or empirical examples of how the “Initial Era” affected individual disciplines or researchers. A more detailed exploration of specific instances where the use of initials had significant consequences would make the findings more tangible. Incorporating case studies or anecdotes from the history of science that illustrate the real-world impacts of the trends identified in the data would enrich the paper. These examples could help ground the analysis in practical outcomes and demonstrate the relevance of the “Initial Era” to contemporary debates.

– Half-baked comparative analysis

Although the paper presents interesting data about different countries and disciplines, the comparative analysis between these groups could be further developed. For example, the reasons behind the differences in initial use between countries with different writing systems or academic cultures are not fully explored. A more in-depth comparative analysis that explains the cultural, linguistic, or institutional factors driving the observed differences in initial use would add nuance to the findings. This could involve a more detailed discussion of how non-Roman writing systems influence name formatting or how specific national academic policies shape author metadata.

– Limited discussion of alternative explanations

While the authors link the decline of the “Initial Era” to technological advancements, other potential explanations, such as changing editorial policies (“technological harmonisation”), shifts in academic prestige, or the influence of global collaboration, are not fully explored. The paper could benefit from a broader discussion of these factors. Expanding the discussion to include alternative explanations for the decline of initial use, and how these might interact with technological changes, would provide a more comprehensive view. Engaging with literature on academic publishing practices, editorial decisions, and global research trends could help contextualize the findings within a wider framework.

Conclusion

This manuscript offers a novel and insightful analysis of the evolution of name usage in academic publications, providing valuable contributions to the fields of bibliometrics, science studies, and research culture. With improvements in clarity, comparative analysis, and the incorporation of case studies, this paper has the potential to make a significant impact on our understanding of how metadata reflects broader societal and technological changes in academia. The authors are encouraged to refine their discussion and expand on the implications of their findings to make the manuscript more accessible and applicable to a wider audience.

Competing interests: None.

Peer review 3

Anonymous reviewer

DOI: 10.70744/MetaROR.15.1.rv3

I started reading this paper with great interest, which flagged over time. As someone with extensive experience both publishing peer-reviewed research articles and working with publication data (Web of Science, Scopus, PubMed, PubMedCentral) I understand there are vagaries in the data because of how and when it was collected, and when certain policies and processes were implemented. For example, as an author starting in the late 1980s, we were instructed by the journal “guide to authors” to use only initials. My early papers were all only using initials. This changed in the mid-late 1990s. Another example, when working with NIH publications data, one knows dates like 1946 (how far back MedLine data go), 1996 (when PubMed was launched), and 2000 (when PubMedCentral was launched) and 2008 (when NIH Open Access policy enacted). There are also intermediate dates for changes in curation policy…. that underlie a transition from initials to full name in the biomedical literature.

I realize that the study covers all research disciplines, but still I am surprised that the authors of this paper don’t start with an examination of the policies underlying publications data, and only get to this at the end of a fairly torturous study.

As a reader, this reviewer felt pulled all over the place in this article and increasingly frustrated that this is a paper that explores the Dimensions database vagaries only and not really the core overall challenges of bibliometric data, irrespective of data source. Dimensions ingests data from multiple sources — so any analysis of its contents needs to examine those sources first.

A few specific comments:

  • The “history of science” portion of the paper focuses on English learned societies in the 17th century. There were many other learned societies across Europe, and also “papers” (books, treatises) from long before the 17th century in Middle-eastern and Asian countries (e.g, see history of mathematics, engineering, governance and policy, etc.). These other histories were not acknowledged by the authors. Research didn’t just spring full-formed out of Zeus’ head.

  • It is unclear throughout if the authors are referring to science, research, which disciplines are or are not included. The first chart on discipinary coverage is Fig 13 and goes back to 1940ish. Also, which languages are included in the analysis? For example, Figure 2 says “academic output” but from which academies? What countries? What languages? Disciplines? Also, in Figure 2, this reviewer would have like to see discussion about the variability in the noisiness of the data over time.

  • The inclusion of gender in the paper misses the mark for this reviewer. When dealing with initials, how can one identify gender? And when working in times/societies where women had to hide their identity to be published…. how can a name-based analysis of gender be applied? If this paper remains a study of the “initial era”, this reviewer recommends removing the gender analysis.

  • Reference needed for “It is just as important to see ourselves reflected in the outputs of the research careers…” (section B).

  • Reference needed for “This period marked the emergence of “Big Science” (Section B). How do we know this is Big Science? What is the relationship with the nature of science careers? Here it would be useful perhaps to mention that postdocs were virtually unheard of before Sputnik.

  • Fig 3. This would be more effective as a % total papers than absolute #.

  • Gradual Evolution of the Scholarly Record. This reviewer would like to see proportion of papers without authors. A lot of history of science research is available for this period, and a few references here would be welcome, as well as a by-country analysis (or acknowledgement that the data are largely from Europe and/or English-speaking countries).

  • Accelerated Changes in Recent Times. Again, this reviewer would like to see reference to scholarship on the history of science. One of the things happening in the post WW2 timeframe is the increase in government spending (in the US particularly) on R&D and academic research. So, is the academy changing or is it responding to “market forces”.

  • Reflective richness of data. “Evolution of the research community” is not described in the text, not is collaborative networks.

  • In the following paragraph, one could argue that evaluation was a driver of change, not a response to it. This reviewer would like to see references here.

  • II. Methodology. (i) 2nd sentence missing “to” “… and full form to refer to an author name…”. (ii) 2nd para the authors talk about epochs, but the data could be (are) discontinuous because of (a) curation policy, (b) curation technology, (c) data sources (e.g., Medline rolled out in the 1960s and back-populated to 1946). (iii) 4th para referes to Figs 3 and 4 showing a marked change between 1940 and 1950, but Fig 3 goes back only to 1960, and Fig 4 is so compressed it is hard to see anything in that time range. (iv) Para 7. “the active publishing community is a reasonable proxy for the global research population”. We need a reference here and more analysis. Is this Europe? English language? Which disciplines? All academia? Dimensions data? (v) Para 12 “In exploring the issue of gender…” see comments above. Gender is an important consideration but is out of scope, in this reviewer’s opinion, for this paper focused on use of initials vs. full name.

  • Listing 1. Is there a resolvable URL/DOI for this query?

  • Figs 9-11, 14, 15. This reviewer would like to see a more fulsome examination / discussion of data discontinuities. Particularly around ~1985-2000.

Discussion

  • The country-level discussion suggests the data (publications included) are only those that have been translated into English. Please clarify. Also, please add references in this section. There are a lot of bold statements, such as “A characteristic of these countries was the establishment of strong national academies.” Is this different from other places in the world? How? In the para before this statement, there is a phrase “picking out Slavonic stages” that is not clear to this reviewer.

  • The authors seem to get ahead of themselves talking about “formal” and “informal” in relation to whether initials or full names are used. And then discuss the “Power Distance” and end up arguing that it isn’t formal/informal … but rather publisher policies and curation practices driving the initial era and its end.

  • And then the authors come full circle on research articles being a technology, akin to a contract. Which is neat and useful. But all the intermediate data analysis is focused on the Dimensions data base and this reviewer would argue should be a part of the database documentation rather than a scholarly article.

  • This reviewer would prefer this paper be focused much more tightly on how publishing technology can and has driven the sociology of science. Dig more into the E. Journal Analysis and F. Technological analysis. Stick with what you have deep data for, and provide us readers with a practical and useful paper that maybe, just maybe, publishers will read and be incentivized to up their game with respect to adoption of “new” technologies like ORCID, DOIs for data, etc. Because these papers are not just expositions on a disciplinary discourse, they are also a window into how science (research) works and is done.

Competing interests: None.

Author response

DOI: 10.70744/MetaROR.15.1.ar

Dear Editor,

Thank you for the editorial assessment of our manuscript “The Rise and Fall of the Initial Era,” and for arranging three substantive peer reviews. We are grateful to all three reviewers for their careful engagement with the paper. We submit a revised manuscript and detailed responses to each reviewer alongside this letter.

We have given particular attention to the three weaknesses identified in your editorial assessment: the clarity, readability, and focus of the introduction and background sections; the depth of the discussion; and the treatment of changes over time in policies, technologies, and data sources. Where reviewer comments point to issues of substance — including all three of these areas — we have engaged in detail and revised the manuscript accordingly. Where comments concern the paper’s underlying scope, framing, or argumentative structure, we have explained our reasoning for retaining the chosen approach while addressing the substantive concerns within it.

The principal revisions are the following.

Introduction and background (clarity, readability, focus). We have restructured the abstract to make the paper’s purpose, contribution, and findings explicit, while keeping its length essentially unchanged. We have added a new closing paragraph to the introduction that articulates a threefold contribution: identification of the Initial Era, multi-layered analysis of its rise and fall, and use of the era as a case study in bibliometric archaeology. The latter concept — which Reviewer 1 rightly noted appeared in the introduction and discussion without connection — has been promoted to a framing thread that runs from the abstract through to the discussion. We have reordered the three “key insights” in Section I.B so that they frame the analysis to come rather than appearing to be drawn from the preceding historical material, addressing Reviewer 1’s specific concern that the third insight did not follow from what preceded it. We have also added a paragraph in Section I.A acknowledging the Eurocentric framing of the historical material and explaining why the seventeenth-century European tradition is the appropriate frame for this paper given its unit of study.

Discussion section (depth and academic context). We have added a new subsection (IV.C, “Alternative drivers and their interaction with technology”) that treats editorial policy, academic prestige economics, and globalisation as candidate explanations alongside the technological account, and have revised the opening of IV.B so that the three factors we ultimately propose are presented as the most plausible synthesis among multiple candidates rather than as the unique explanation. Reviewer 2’s comment on alternative explanations was particularly valuable in shaping this revision.

Policies, technologies, and data sources over time. Reviewer 3’s central substantive concern — that the paper treats Dimensions as a transparent window onto the scholarly record rather than as a layered artefact — was well-taken and is reinforced by the editorial assessment. We have added a new methodology subsection (II.A, “The layered provenance of bibliographic data”) that traces the policy and technology timeline of the major data sources Dimensions ingests, including the dates that create coverage and metadata discontinuities. The new subsection explicitly signposts Section III.F (Technological Analysis) as the place where the paper engages with these effects analytically, and the discussion subsection IV.D (formerly IV.C, on bibliometric archaeology) returns to the methodological implications.

Gender analysis. We note that the reviewers disagreed on this section, with Reviewer 2 considering it a strength and Reviewer 3 recommending its removal. With the editorial assessment leaving the question open, and given that we believe Reviewer 3’s critique points to a real weakness in our framing rather than to an irreducible flaw in the analysis, we have substantially strengthened the framing and methodological caveats of Section III.G rather than removing it. The section is now explicitly framed as an analysis of gender visibility in the scholarly metadata rather than of gender participation in research, and we have articulated the natural-experiment logic of the 2002 PubMed metadata transition more carefully. We have named the Lockhart et al. critique of name-to-gender attribution methods explicitly and acknowledged its implications for our work.

New references. We have added four references to address specific gaps identified by the reviewers: Weinberg (1961) and de Solla Price (1963) for the “Big Science” claim; Stephan (2012) for the post-war restructuring of research careers, including the post-Sputnik R&D context; and Hofstra et al. (2020) for the diversity-and-research-output claim. A footnote refers readers to the Wilsdon Review of metric-driven evaluation in the new alternative-drivers subsection.

Items we have considered and declined. For transparency, we note the reviewer suggestions we have considered carefully and ultimately declined, with reasons given in the individual responses: removal of the gender analysis (Reviewer 3); presentation of Figure 3 as a percentage rather than absolute counts (Reviewer 3); revision of the colour palette in Figure 8 (Reviewer 1, on accessibility grounds); addition of further case studies beyond the three journal-level cases already in Section III.E (Reviewer 2); and restructuring the paper to begin from the policy and curation history of bibliographic data (Reviewer 3). We have also retained the abstract at approximately its original length while restructuring it for clarity, rather than expanding it to 250 words as Reviewer 1 suggested.

The net length increase from these revisions is approximately 800 words. We have tried to ensure that the additions strengthen the paper’s central argument rather than diluting it.

We are grateful for the rigour of the review process and for the constructive nature of all three reviews.

With best regards,

Simon J Porter and Daniel W Hook

Response to Reviewer 1 (Dmitry Kochetkov)

We thank Dr Kochetkov for a careful, constructive review and for the positive overall assessment. We have addressed each of his points below. References to numbered changes (e.g. “Change 7”) refer to the consolidated change list at the end of this document.

R1.1 — Expand the abstract to approximately 250 words. We have restructured the abstract to make purpose, contribution, and findings explicit, but have kept the length essentially unchanged. We agree with the reviewer that the original abstract did not convey the paper’s contribution clearly, but we judge that the issue was structural rather than one of length, and that a 250-word abstract risks duplicating the introduction. The revised abstract now states explicitly what the paper does, what data it uses, what is argued, and what is implied [Change 1].

R1.2 — The introduction starts with very broad context whose connection to the research object is not immediately clear; few references make it difficult to determine whether the authors are citing others or their own findings. We have added a closing paragraph to the introduction that states the paper’s purpose, objectives, and threefold contribution explicitly [Change 2]. We have also added references in two places where the reviewer (and Reviewer 3) noted that claims were unsupported: for “Big Science” we now cite Weinberg (1961) and de Solla Price (1963), and for the post-war restructuring of research careers we cite Stephan (2012) [Change 4]. For the diversity-and-research-output claim we now cite Hofstra et al. (2020) [Change 5].

R1.3 — The transition to the main topic is not well-defined, and “bibliometric archaeology” appears once in the introduction and once in the discussion without connection. We have promoted bibliometric archaeology to a more prominent framing concept. It now appears in the abstract, is named explicitly as the third element of the paper’s contribution in the new closing paragraph of the introduction, is referenced in the new methodology subsection on the layered provenance of bibliographic data, and is returned to in the Discussion. We hope the reader now experiences this as a coherent framing thread rather than two isolated mentions [Changes 1, 2, 7].

R1.4 — State the purpose and objectives clearly in both the abstract and the introduction. Addressed in Changes 1 and 2.

R1.5 — Elaborate on the contribution of the study in the introduction. The new closing paragraph of the introduction articulates a threefold contribution: identification of the Initial Era, multi-layered analysis of its rise and fall, and use of the era as a case study in bibliometric archaeology [Change 2].

R1.6 — The same applies to the background section: broad context, weak connection. We have made two changes here. First, we have reordered the three “key insights” (Gradual Evolution, Accelerated Changes, Reflective Richness) so that they now precede rather than follow the discussion of Figs 2 and 3, framing the analysis to come rather than emerging from the historical material — this directly addresses the reviewer’s specific concern that the third insight does not follow from the preceding analysis [Change 6]. Second, we have added linking sentences and references throughout the background to connect the historical material to the research object more tightly [Changes 4, 5].

R1.7 — Page 4: the conclusions from a literature review do not all follow from previous analysis; specifically, “Reflective Richness of Data” does not. This is the same concern as R1.6 and is fully addressed by the reordering in Change 6. The three insights now frame the data presentation rather than being presented as conclusions from it.

R1.8 — Shorten the introduction and literature review; the introduction text could be a standalone contribution. We have undertaken targeted tightening of the introduction (the reordering in Change 6 and the explicit-purpose paragraph in Change 2) but have retained the substantive content. The literary parallels (Austen, Conan Doyle) and the contextual material on pseudonymity and identity that the reviewer flags as potentially extraneous serve, in our view, to motivate the central argument that name-form conventions reflect deeper questions about visibility and identity in scholarly communication — a theme the paper develops empirically. We are persuaded by the reviewer (and by the editorial assessment) that the introduction needed structural improvement; we are less persuaded that significant content cuts would strengthen the paper’s argument.

R1.9 — In the methodology, justify the structure of presenting the results. The new closing paragraph of the introduction now signposts the analytical structure (paper-level, author-level, country, discipline, journal, technology, gender), which we hope gives the reader the orientation requested [Change 2].

R1.10 — Provide a footnote with a link to the code and dataset early in the methodology. The data and code are referenced in the existing Data Availability statement, which contains a Figshare DOI providing access to all data, all code, and a complete reproduction package. We have not duplicated this in the methodology section as it would risk redundancy, but we are happy to add a forward-pointer if the editor prefers.

R1.11 — In the discussion, place results more clearly in the academic context, ideally with references from the introduction reappearing. We have substantially revised the discussion. A new subsection (IV.C, “Alternative drivers and their interaction with technology”) engages with editorial policy, prestige economics, and globalisation as alternative or complementary explanations to the technological account, with references to existing and newly-introduced literature [Change 12]. We have also restructured the opening of IV.B to acknowledge that the technological account is one perspective among several [Change 11], and have revised the gender material in IV.B to align with the strengthened framing of Section III.G [Change 13c].

R1.12 — Discussion C is interesting but seems more related to the introduction than the results; it would benefit from a more thorough justification. The new bibliometric-archaeology framing thread (R1.3 above) means that the Discussion section on the future of bibliometric archaeology — now IV.D following the insertion of IV.C — connects directly back to a concept introduced in the abstract, the introduction, and the methodology. We hope this answers the reviewer’s concern that the section was inadequately connected to the rest of the paper.

Image remarks 1 — The data source for the images should be specified in the background. The new methodology subsection on the layered provenance of bibliographic data (II.A) makes the Dimensions basis of all figures explicit early in the paper [Change 7]. We have not added the data source to each figure caption, since this would create substantial visual clutter and the source is now stated very prominently both in the background discussion and the methodology.

Image remarks 2 — The colour distinction between China and Russia in Figure 8 is not very clear. We have considered this carefully. The current colour palette has been chosen to maintain accessibility for partially-sighted readers, and we have found it difficult to improve the China/Russia distinction without compromising accessibility for other categories. We have therefore retained the current palette, but note that the reviewer’s concern is registered and we will look at this again in case we submit to a journal and need to make changes at a proof stage and a better solution presents itself.

Image remarks 3 — Grey lines in Figs 9–11 make figures difficult to read; their meaning is not indicated in the legends of Figs 10 and 11. We have updated the captions of Figs 10 and 11 to match Fig 9 in stating explicitly that the grey dotted lines represent all countries that consistently participate in more than 500 papers per year, plotted for context [Change 10].

We thank the reviewer again for the constructive engagement.

Response to Reviewer 2 (Erjia Yan)

We thank Professor Yan for a thorough and constructive review and for the positive assessment of the paper’s strengths, including the gender analysis (which we note has been the subject of disagreement between reviewers and which we discuss further below). We address each weakness in turn.

R2.1 — Lack of clarity and readability; condense the introduction, remove unrelated historical accounts, and lead the audience to key objectives earlier. We have made two structural improvements to the introduction. First, we have added a new closing paragraph that articulates the paper’s threefold contribution explicitly, so that the reader reaches a clear statement of purpose before the methodology section [Change 2]. Second, we have reordered the three “key insights” so that they frame the analysis to come rather than appearing to be drawn from the preceding historical discussion [Change 6]. We have also added a paragraph in Section I.A acknowledging the Eurocentric framing of the historical material and explaining why the seventeenth-century European tradition is the appropriate frame for this paper given its subject [Change 3]; this also helps clarify what the paper is and is not about.

We have not undertaken substantial cuts to the historical material. The literary parallels and contextual discussion of pseudonymity, while not strictly necessary to the data analysis, serve to motivate the broader argument that author-name conventions are an instrument of visibility and identity in the scholarly record. We accept the reviewer’s point about clarity — and the editorial assessment’s reinforcement of it — but in our view the issue is structural rather than one of excess content, and we have addressed it through reorganisation rather than deletion.

R2.2 — Missing empirical case studies; incorporate concrete examples or case studies from the history of science. The paper already contains case studies of three named journals — Journal of Biological Chemistry, Tetrahedron, and The BMJ — in Section III.E, where we trace the policy and platform events that drove their distinct trajectories through the Initial Era and beyond. These are concrete cases that ground the macro-pattern, and the post-1995 Journal of Biological Chemistry transition coinciding with the Highwire launch is a particularly clean instance of the technology-driven mechanism we propose. We hope that on a second reading the reviewer will recognise these as the case-study material the review calls for.

We have considered adding additional case studies — for example a named individual researcher whose contributions were obscured by initial-form publication — but have decided against doing so. The paper is already substantial in length, and the editorial assessment specifically flags clarity and focus as concerns; further extension would work against that goal. We have, however, made the connection between Section III.E and the broader argument more explicit by referencing the journal-level cases in the new alternative-drivers subsection [Change 12].

R2.3 — Half-baked comparative analysis; the reasons behind differences between countries with different writing systems or academic cultures are not fully explored. We have added a paragraph acknowledging the Eurocentric framing of the historical material, which goes some way toward addressing the writing-systems point — making explicit that what we study is the modern globalised journal article as a particular technology, with the conventions it inherited from a particular tradition [Change 3].

The deeper comparative analysis the reviewer suggests — disentangling cultural, linguistic, institutional, and policy effects on a country-by-country basis — would constitute a substantial separate study. The current paper is limited to the observation that these effects are visible in the data and that they appear, on the evidence of Fig. 12 and Section III.C, to be only weakly correlated with measures of cultural formality such as Hofstede’s Power Distance. We have taken the more cautious view that a deeper comparative analysis is beyond the scope of the present paper, and have made this explicit in the existing discussion. We are grateful to the reviewer for identifying this as a productive direction for further work.

R2.4 — Limited discussion of alternative explanations beyond technology — editorial policies, academic prestige, global collaboration are not fully explored. This is, we believe, the reviewer’s most substantive point and we have engaged with it directly. We have added a new subsection IV.C (“Alternative drivers and their interaction with technology”) that treats editorial policy and house style, academic prestige economics, and globalisation as candidate explanations alongside the technological account [Change 12]. We have also revised the opening of IV.B so that the three factors we ultimately propose (technological harmonisation, globalisation, trust) are presented as the most plausible synthesis among multiple candidates rather than as the unique explanation [Change 11]. The new IV.C material engages with the Pardo-Guerra evaluation literature already cited [22] and refers, in a footnote, to the Wilsdon Review’s critique of metric-driven evaluation.

We thank the reviewer for the encouraging conclusion and hope that the revised discussion now provides the comprehensive view requested.

Response to Reviewer 3

We thank the reviewer for a detailed and challenging review. The reviewer’s critique has prompted substantial revisions to the paper, particularly around the policy and curation history of bibliographic data, the framing of the historical material, and the methodological caveats around the gender analysis. We address each comment below.

On the overall framing. The reviewer argues that the paper should have begun with an examination of the policies underlying publications data, rather than reaching this material at the end of “a fairly torturous study.” We accept the substantive point that policy and curation history deserve more prominence, and we have responded by adding a new methodology subsection (II.A) on the layered provenance of bibliographic data, which traces the policy and technology timeline of the major data sources Dimensions ingests — including the dates the reviewer specifically cites (1946 MedLine retrospective coverage, 1996 PubMed launch, 2000 PubMedCentral, 2002 PubMed redesign) [Change 7]. The new subsection explicitly acknowledges that some of the discontinuities in our data are features of the recording systems rather than of scholarly practice itself, and it signposts Section III.F as the place where the paper engages with this distinction analytically.

We have not, however, restructured the paper so that this material is its starting point. Our reasoning is that the paper has two contributions — identifying the Initial Era as a phenomenon, and showing how it is illuminated by reading bibliographic data as a layered artefact — and beginning with the second risks losing the first. We hope the reviewer will accept that the new II.A places the policy/curation context where it belongs methodologically, and that Section III.F now has the prominence it deserves through the explicit signposting in the methodology and the new discussion subsection IV.C.

Specific comments:

R3.1 — The history-of-science portion focuses on English learned societies in the seventeenth century and ignores other scholarly traditions. This is a fair criticism and we have addressed it directly. A new paragraph in Section I.A acknowledges that scholarly traditions of substantial sophistication existed long before, and in parallel with, seventeenth-century European learned societies — in mathematics, astronomy, medicine, and other areas across the Islamic world, China, India, and elsewhere — and explains that our focus on the Philosophical Transactions and its successors reflects not a claim about the origins of scholarship but a narrower observation about the modern journal article as a specific technology with particular conventions [Change 3]. We are grateful to the reviewer for highlighting an oversight that the paper now corrects.

R3.2 — Unclear which disciplines, languages, and academies are covered; Fig 2 says “academic output” but from where? The new methodology subsection II.A discusses the data sources directly [Change 7], and we have added a sentence to the existing epoch discussion that draws attention to the standard-deviation envelope in Fig 2 as a visual indicator of the small n in the early period [Change 8]. The figure now serves as its own data-availability commentary. The paper’s scope is determined by Dimensions’ coverage and is necessarily skewed toward English-language, Western-curated sources for the earlier periods; the paper makes this explicit in the methodology section.

R3.3 — The inclusion of gender misses the mark; with initials, how can one identify gender? Recommend removing. We have engaged with this point at length. We note that Reviewer 2 considers the gender analysis to be a strength of the paper, and that the editorial assessment leaves the question open. With that in mind — and given that we believe the reviewer’s critique points to a real weakness in our framing rather than to an irreducible flaw in the analysis — we have substantially strengthened the framing and methodological caveats of the gender section rather than removing it.

The reviewer’s central concern is correct: gender cannot be inferred from initials. But this is precisely the finding the analysis is designed to illuminate, not a fatal objection to it. The gender analysis is not a measurement of gender participation in research; it is a measurement of gender visibility in the scholarly metadata. The pre-2002 initial-form convention rendered a very large fraction of authors gender-indeterminate by name; the 2002 PubMed metadata change made gender suddenly inferrable for a much larger share of the population. The relative jump sizes for female-coded and male-coded names following this metadata change cannot be a participation effect — participation does not change discontinuously in a single year — and the disproportionate steepness of the female-coded rise is therefore most plausibly read as a visibility effect: the initial-form convention was hiding women, in the metadata, more than it was hiding men. This is a structural argument at the population level and makes no claims about individual researchers.

We have rewritten the opening of Section III.G to make this framing explicit and to acknowledge the Lockhart et al. critique of name-to-gender attribution methods [Change 13a]. We have rewritten the closing paragraph of the section to articulate the natural-experiment logic of the 2002 transition explicitly [Change 13b]. We have also revised the corresponding sentence in Section IV.B to make the claim more careful — “render systematically less visible in the scholarly metadata” rather than “hide and marginalise” [Change 13c]. We hope the reviewer will agree that, with these revisions, the analysis now defends itself against the methodological objection raised.

R3.4 — Reference needed for “It is just as important to see ourselves reflected in the outputs of the research careers…” Added [Change 5]; we now cite Hofstra et al. (2020) on the diversity-innovation relationship.

R3.5 — Reference needed for “Big Science”; mention that postdocs were virtually unheard of before Sputnik. Added [Change 4]; we now cite Weinberg (1961) for the term itself, de Solla Price (1963) for the bibliometric framing, and Stephan (2012) for the postdoctoral and R&D-spending context the reviewer specifically mentions.

R3.6 — Fig 3 would be more effective as a percentage of total papers than absolute numbers. We have considered this and have decided to retain the absolute-count presentation. Our reasoning is that Fig 3’s purpose is to convey the emergence of large-collaboration science as a new phenomenon in the late 1970s — papers with 500 or 1000 co-authors — and the absolute-count presentation conveys this directly: such papers did not exist before. A percentage presentation would render these papers as small fractions of the (much larger) total publication output, obscuring the rise rather than highlighting it. We acknowledge that the reviewer’s preferred presentation would be informative for a different question and would make a useful supplementary analysis, but we believe the question Fig 3 is designed to answer is best served by the current presentation.

R3.7 — “Gradual Evolution of the Scholarly Record”: would like to see proportion of papers without authors, history-of-science references, by-country analysis or acknowledgement of geographic skew. The new Section I.A paragraph acknowledges the geographic skew of the historical material directly [Change 3]. The new methodology subsection II.A discusses the data-coverage limitations of the early period explicitly [Change 7]. A full analysis of papers without identified authors would constitute a separate paper and is beyond the scope of the present study; we note in passing that the early Philosophical Transactions figures we already report (171 of 357 articles in 1665–1669 with identifiable authors) implicitly provide the by-journal version of this for the seventeenth century.

R3.8 — “Accelerated Changes in Recent Times”: reference to scholarship on the history of science; post-WW2 increase in government R&D spending. Addressed by the new Stephan (2012) reference [Change 4], which covers the post-Sputnik R&D context the reviewer suggests, and by the new alternative-drivers subsection IV.C, which engages with the Pardo-Guerra (2022) evaluation literature [Change 12].

R3.9 — “Reflective Richness of Data”: evolution of research community and collaborative networks not described. This concern is partially addressed by the reordering of the three insights [Change 6] — they now frame what is to come, rather than appearing to summarise what has been shown. The “richness of data” point is then exemplified by the analyses that follow.

R3.10 — “Evaluation as a driver” needs references. Pardo-Guerra (2022) is already cited as reference [22]; the new subsection IV.C draws on this more explicitly [Change 12], and a footnote points to the Wilsdon Review of metric-driven evaluation.

R3.11 — Methodology specific points. (i) Missing “to” — fixed [Change 9]. (ii) Discontinuities from curation policy/technology/sources — addressed by new subsection II.A [Change 7]. (iii) Fig 3 dates — Fig 3 is restricted to 1960 onwards by deliberate choice (papers with 100+ authors do not exist before this); the marked change between 1940 and 1950 is most clearly seen in Fig 4, and we have added the variance observation noted in [Change 8] to clarify Fig 2 in the same period. (iv) “Active publishing community as proxy” — we have not added a specific reference here as the claim is qualified in the text and the limitations are explicitly acknowledged. (v) Gender out of scope — addressed at length in our response to R3.3.

R3.12 — Listing 1: is there a resolvable URL/DOI for this query? The Listing 1 query, along with all derived analyses, is available in the Figshare repository referenced in the Data Availability statement. We have not added a separate DOI for the listing alone, since it is a component of the larger reproduction package.

R3.13 — Figs 9–11, 14, 15: more fulsome examination of data discontinuities, particularly around 1985–2000. The new subsection II.A engages with the systemic discontinuities directly [Change 7]; Section III.F continues to treat the 2002 PubMed transition explicitly. The 1985–2000 period is one of multiple overlapping changes (the introduction of online publishing, journal-level policy revisions, the launch of PubMed and Crossref) and we address it through the journal-level cases in Section III.E and the technological analysis in Section III.F. A more detailed disentangling of these effects across multiple disciplines and countries is, again, a productive direction for separate work.

Discussion comments — country-level data appear to be only English; references missing; “strong national academies” undefended; “picking out Slavonic stages” unclear; the paper appears to come full circle on research articles as technology. The “Slavonic stages” wording was a stylistic choice we no longer endorse — we have changed it to “Slavonic-language countries” [Change 15]. We acknowledge that the national-academies discussion is somewhat speculative, and we have re-read the paragraph to ensure it now reads clearly as a conjecture rather than as an established claim; the paragraph already opens with the statement that “isolating the effect of national academies in research culture agenda-setting is beyond the scope of the current work.” On the question of formality versus publisher policy and curation: the reviewer’s reading is correct that the paper argues, on the evidence, that publisher policies and curation practices were more important drivers than cultural formality at the country level. This is, in our view, a finding rather than a contradiction. The paper’s central claim is that the scholarly article is a technology shaped by multiple interacting forces; we identify those forces empirically and conclude that — for the Initial Era specifically — the technological and policy layer was the most consequential.

We thank the reviewer again for the rigour of the engagement, which has materially improved the paper.

Consolidated List of Changes to the Manuscript

All changes below are referenced by number in the responses to individual reviewers. The net length increase is approximately 800 words. New bibliography entries appear at the end.

Change 1 — Abstract restructured

Section: Abstract. Addresses: R1.1, R1.4, R1.5; editorial assessment (clarity/focus). Effect: Restructured to make purpose, contribution, and findings explicit. Length essentially unchanged.

Change 2 — New closing paragraph in introduction

Section: Introduction (after the “The paper is organised” paragraph). Addresses: R1.4, R1.5, R1.6; editorial assessment. Effect: New paragraph stating threefold contribution: identification of the Initial Era; multi-layered analysis; bibliometric archaeology framing.

Change 3 — Eurocentrism acknowledgement

Section: I.A. Addresses: R3.1; editorial assessment. Effect: New paragraph acknowledging that scholarly traditions existed before and in parallel with seventeenth-century European learned societies, and explaining why the journal article as a specific technology is nonetheless the appropriate frame for this paper.

Change 4 — “Big Science” references added

Section: I.B (the “emergence of ‘Big Science’” paragraph). Addresses: R1.2, R3.5, R3.8. Effect: Added Weinberg (1961) for the coinage of “Big Science”, de Solla Price (1963) for the bibliometric framing, and Stephan (2012) for the post-Sputnik R&D and postdoc context.

Change 5 — Diversity reference added

Section: I.B (the “see ourselves reflected” sentence). Addresses: R3.4. Effect: Added Hofstra et al. (2020) on the diversity–innovation relationship.

Change 6 — Three insights reordered

Section: I.B. Addresses: R1.6, R1.7, R3.9. Effect: The numbered list of three insights (Gradual Evolution, Accelerated Changes, Reflective Richness) has been moved to before the discussion of Figs 2 and 3, with the lead-in changed from “From the overarching observations and detailed examples discussed, three key insights emerge…” to “Three insights frame the analysis that follows…”. The insights now frame the analysis rather than appearing to be drawn from preceding material.

Change 7 — New methodology subsection II.A

Section: Methodology (new subsection “The layered provenance of bibliographic data” inserted at start of Section II). Addresses: R3 (overall framing), R3.7, R3.11, R3.13; editorial assessment. Effect: New subsection traces the policy/technology timeline of major data sources Dimensions ingests (MedLine 1946 retrospective and 1960s rollout, PubMed 1996, PubMedCentral 2000, PubMed 2002 redesign, Crossref DOI infrastructure, JATS XML). Acknowledges that some discontinuities are features of recording systems rather than of scholarly practice. Signposts Section III.F as the analytical engagement with this distinction. New label sec:provenance added.

Change 8 — Fig 2 variance observation

Section: Methodology (existing epoch discussion paragraph). Addresses: R3.2, R3.11. Effect: Added a sentence drawing attention to the width of the standard-deviation envelope in Fig 2 as a visual indicator of small n in the anecdotal epoch, reinforcing the argument for the 1945 cutoff.

Change 9 — Missing “to” corrected

Section: Methodology, second sentence. Addresses: R3.11(i). Effect: “full form to refer an author name” corrected to “full form to refer to an author name”.

Change 10 — Captions of Figs 10 and 11 clarified

Section: Results, geographic analysis figures. Addresses: R1 image remarks 3. Effect: Captions of Figs 10 and 11 now state explicitly that the grey dotted lines represent all countries that consistently participate in more than 500 papers per year, plotted for context, matching Fig 9’s caption.

Change 11 — Opening of IV.B restructured

Section: Discussion, IV.B (“The fall of the Initial Era”). Addresses: R2.4, R1.11. Effect: Opening sentence — “The fall of the Initial Era is fairly clear cut from a technological perspective” — replaced with text that explicitly acknowledges multiple plausible accounts and presents the technological account as one perspective. The three factors that follow are now framed as the most plausible synthesis among candidates rather than the unique explanation.

Change 12 — New discussion subsection IV.C

Section: Discussion (new subsection “Alternative drivers and their interaction with technology” inserted between the previous IV.B and the existing IV.C, which becomes IV.D). Addresses: R2.4, R1.11; editorial assessment (deepen discussion). Effect: New subsection treating editorial policy/house style, academic prestige economics, and globalisation as alternative or complementary explanations to the technological account. Includes a footnote referencing Wilsdon et al., The Metric Tide (HEFCE, 2015). Cross-references the existing journal cases in III.E (new label sec:journal) and the Technological Analysis section (new label sec:tech).

Change 13 — Gender section reframing

Section: III.G (Gender Analysis) and IV.B. Addresses: R3.3. Effect: Three sub-changes:

13a: new framing paragraph at the start of III.G distinguishing visibility from participation, acknowledging Lockhart et al. (2023) critique of name-to-gender attribution and stating clearly that the analysis makes no claims about specific researchers.

13b: closing paragraph of III.G replaced with a natural-experiment framing of the 2002 PubMed transition, arguing that the steeper proportional rise of female-coded names cannot be a participation effect (which would require a discontinuous demographic shift in a single year) and is therefore most plausibly a visibility effect.

13c: corresponding sentence in IV.B revised from “hide and marginalise the contribution of women to research” to “render the contribution of women to research systematically less visible in the scholarly metadata”, aligning with the strengthened framing in III.G.

Change 14 — New bibliography entries

Section: Bibliography (Initials.bib). Addresses: Changes 4 and 5. Effect: Four new entries appended:

  • weinberg_impact_1961 — A. M. Weinberg, “Impact of Large-Scale Science on the United States,” Science 134(3473), 161–164 (1961). DOI: 10.1126/science.134.3473.161
  • desolla_little_1963 — D. J. de Solla Price, Little Science, Big Science. Columbia University Press, New York (1963). DOI: 10.7312/pric91844
  • stephan_economics_2012 — P. Stephan, How Economics Shapes Science. Harvard University Press, Cambridge, MA (2012). DOI: 10.4159/harvard.9780674062757
  • hofstra_diversity_2020 — B. Hofstra, V. V. Kulkarni, S. Munoz-Najar Galvez, B. He, D. Jurafsky, and D. A. McFarland, “The Diversity–Innovation Paradox in Science,” Proceedings of the National Academy of Sciences 117(17), 9284–9291 (2020). DOI: 10.1073/pnas.1915378117
Change 15 — “Slavonic stages” wording corrected

Section: Geographical analysis (introduction to Fig 11). Addresses: R3 (Discussion comments). Effect: “picking out Slavonic stages” corrected to “picking out Slavonic-language countries”.

Items considered and declined

Removal of the gender analysis (R3): addressed by strengthening the framing and methodological caveats (Change 13) rather than removal, given Reviewer 2’s contrary view and the editorial assessment leaving the question open.

Fig 3 as percentage rather than absolute counts (R3.6): declined; absolute counts better convey the emergence of large-collaboration science as a new phenomenon. Acknowledged as a useful supplementary analysis.

Fig 8 colour palette revision (R1 image remarks 2): declined on accessibility grounds; current palette has been chosen for partially-sighted readers.

Additional case studies (R2.2): declined; existing journal cases in Section III.E provide the case-study material, and the editorial assessment’s emphasis on clarity and focus weighs against further extension.

Restructure paper to start from policy/curation history (R3): declined; addressed instead through new subsection II.A (Change 7), which places the policy context where it belongs methodologically while preserving the paper’s argumentative structure.

Abstract expansion to ~250 words (R1.1): declined; the issue identified is structural and addressed by restructuring (Change 1) rather than by adding length.

Leave a comment