Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

A data-driven comparative analysis on citation intents across scientific disciplines

  • Tuan Anh Phan,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft

    Affiliation Department of Computer Engineering, Chung-Ang University, Seoul, Republic of Korea

  • Seohyun Nam,

    Roles Data curation, Investigation, Methodology, Resources, Writing – review & editing

    Affiliation Department of Computer Engineering, Chung-Ang University, Seoul, Republic of Korea

  • Jason J. Jung

    Roles Conceptualization, Formal analysis, Funding acquisition, Investigation, Project administration, Supervision, Writing – review & editing

    j3ung@cau.ac.kr, j2jung@gmail.com

    Affiliation Department of Computer Engineering, Chung-Ang University, Seoul, Republic of Korea

Abstract

The goal of this paper is to conduct a first comparative analysis on citation intents across multiple scientific disciplines. To this end, we conducted a data-driven investigation into the characteristics of citation intents utilizing a large-scale dataset of 17,717 academic papers and 2,345,674 citation contexts from 21 Web of Science subject categories. We analyzed five key characteristics: i) inter-category distribution of citation intents, ii) average number of citation intents per article, iii) age distribution of citation intents, iv) similarity of the age distributions of citation intent, and v) predictability of the number of citation intents. We found out that there are significant differences in all five characteristics across different research disciplines. Specifically, Philosophy, Sociology, Area, and History prioritize Background and Motivation intents while exhibiting minimal engagement with Uses and Extends. Conversely, Computer Science AI represents a unique case, characterized by a dominant prevalence of Uses and Extends intents alongside the least frequent use of background-oriented citations. Furthermore, the age distributions of citation intents exhibit the lowest similarity in Area, and History. In terms of predictability, Biotechnology and Plant Science exhibit the highest correlation levels in citation intent counts, whereas Sociology, Philosophy, and History yield the lowest correlation coefficients. Our comparative analysis not only uncovers cross-disciplinary variations and establishes a standardized framework for the comparative analysis of citation intents, but also has applications in advancing the normalization of scholarly indicators, facilitating trend detection, and supporting curriculum design for specific scientific disciplines.

Introduction

In the field of bibliometrics, the analysis of citation intents is very important as it reflects the rhetorical functions of citations, representing the specific reasons why authors use the citations to refer to other scholarly works [14]. Consequently, uncovering the characteristics of the citation intents at a disciplinary level is essential to deciphering the unique mechanisms through which different scholarly areas validate and communicate academic knowledge.

However, the existing literature on citation intents in bibliometrics has only predominantly centered on two key directions: i) defining robust taxonomies and establishing benchmark datasets [16]; and ii) developing novel methodologies to enhance automated classification performance [3,4,719]. Consequently, studies exploring the characteristics of citation intent at both disciplinary and cross-disciplinary scales remain remarkably scarce. Driven by this necessity and the aforementioned gap, this study represents a first effort to characterize citation intents across multiple disciplines, aim to investigate whether significant variations exist regarding their fundamental characteristics.

To this end, we first sampled 17,717 papers with 2,345,674 citation contexts from multiple Web of Science (WoS) scienctific disciplines. We also built a citation intent classifier to automatically generate the citation intents for each citation context in our dataset. Finally, we measured and compared set of five objective characteristics such as: i) inter-category distribution of citation intents, ii) average number of citation intents per article, iii) age distribution of citation intents, iv) similarity of the age distribution of citation intents, and v) predictability of the number of citation intents. Detail explanation of each characteristic is given in Section Materials and Methods. By conducting this comparative analysis, our study not only provides critical insights into cross-disciplinary variations in citation intents but also establishes a standardized framework for analyzing the pattern of citation intent between other scientific entities.

The remainder of this paper is organized as follows. First, we introduce our dataset, the citation intent taxonomy, and the methodology employed to train the classification model. Next, we detail the characteristics used in this study and present our empirical results and findings. Finally, we provide an in-depth discussion of the results and conclude our research.

Materials and methods

Our proposed framework comprises four stages: Data collection, Citation intent classifier construction, Citation intent prediction, and Analysis. The Data collection phase aims to collect highly cited articles from multiple scientific disciplines along with their bibliometric information and citation contexts. The Citation intent classifier construction phase targets to develop a high-quality model to facilitate the automated prediction of citation intents given citation contexts. In the Citation intent prediction stage, citation contexts are fed into our classification model to automatically predict corresponding citation intents. At the final step, we will perform comparative analyses of citation intent characteristics across disciplines. The overview of our framework is illustrated in Fig 1. In the following sections, we will describe the three primary stages: Data collection, Citation intent classifier construction, and Analysis. Regarding the Citation intent classifier construction, we present the taxonomy for citation intent we used in this study and how we train our citation intent classifier.

thumbnail
Fig 1. The overview of our framework in this study.

https://doi.org/10.1371/journal.pone.0358327.g001

Data collection

The data collection process began with the retrieval of information across multiple disciplines. Web of Science (WoS) was selected as the primary data source, as it is widely regarded as one of the most authoritative and comprehensive databases for scholarly publications and citation data [20]. To ensure disciplinary diversity, we selected 21 distinct WoS subject categories (hereafter referred to as disciplines or fields). Within each discipline, the top five Q1 journals were identified to prioritize high-quality research. For each selected journal, the top 5% most-cited papers were identified from Scopus database [21]. After that, review journals were expressly excluded to maintain a strict focus on original research articles. To ensure data robustness, our dataset only included journals with at least 20 collected cited papers; journals falling below this threshold were excluded from the study. For each cited paper, bibliographic resolution was primarily conducted using Scopus metadata [21]. To collect the citing papers and citation contexts for the cited articles, OpenAlex [22] was employed for DOI and title matching. Once the papers were resolved, the Semantic Scholar Graph API [23] was utilized to collect citing papers and their associated contexts. This API facilitated the extraction of rich metadata, including paper identifiers and specific citation contexts essential for our analysis. Our multi-stage corpus collection procedure is shown in Fig 2.

thumbnail
Fig 2. Flowchart of the multi-stage corpus collection procedure.

https://doi.org/10.1371/journal.pone.0358327.g002

In addition, during the final retrieval phase (Stage 6), a fraction of candidate papers could not be fetched due to substantial cross-disciplinary variations in database coverage. To provide a granular view of this data loss, the number of discarded records at this stage is detailed in Table 1. The results indicate that data loss occurred disproportionately across domains. As quantified in Table 1, while fields such as Plant Sciences () and Biotechnology () experienced minimal loss, domains like Film, Radio, TV () and Physics () were far more severely affected a disparity that directly reflects the uneven coverage of indexing services across academic disciplines. In addition, in the final corpus, a sample imbalance remains, as certain disciplines comprise fewer than 100 documents (e.g., History), whereas others contain over 1,000 documents (e.g., Chemistry and Medicine). This stems directly from intrinsic differences in natural publication volume and substantial disparities in database coverage across disciplines. Finally, our corpus comprises 17,717 cited papers and 2,345,674 citation contexts.

thumbnail
Table 1. Full name, abbreviation, retrieval loss rates, and counts of cited articles and citation contexts across disciplines.

https://doi.org/10.1371/journal.pone.0358327.t001

Taxonomy for citation intent

This study utilizes the citation intent framework grounded in the MultiCite dataset [4]. We selected this source due to its accurate, comprehensive, and diverse set of definitions, which provides a highly reliable foundation for characterizing citation purposes across our analysis. The detail definition of the taxonomy for citation intents in Lauscher et al. [4] is given as follows:

  • Background: The cited paper provides foundational context and essential domain-specific information.
  • Motivation: The cited paper provides the rationale for the study, such as the need for specific data, goals, or methods.
  • Uses: The citing paper directly employs ideas, methodologies, tools, or techniques from the cited work.
  • Extends: The citing paper builds upon, adapts, or enhances existing concepts and technical solutions.
  • Similarities: The citing paper identifies shared characteristics or parallel findings between the studies.
  • Differences: The citing paper delineates contrasts, discrepancies, or conflicting results relative to the referenced work.
  • Future work: The cited paper suggests potential avenues and directions for subsequent research exploration.

In the following part, we will present how we use this taxonomy and the MultiCite dataset to build the citation intent classifier.

Citation intent classifier

In this study, to build the classifier for citation intents, we follows a fine-tuning paradigm by first employing a pre-trained language model as our architectural backbone, and then subsequently fine-tuning it on the MultiCite dataset to perform supervised classification. The language model we selected was SciBERT [7] as it is a pre-trained language model trained on a large-scale scientific corpus, ensuring its suitability for processing the citation contexts. The overview of our classifier is presented in Fig 3.

thumbnail
Fig 3. Architecture of our citation intent classifier utilizing the SciBERT backbone.

https://doi.org/10.1371/journal.pone.0358327.g003

The input citation sentence C was first partitioned into a sequence of discrete tokens, denoted as , where N represents the total number of tokens in this context. The list of tokens were then added with two special tokens [CLS] and [SEP] as , and were mapped into a d-dimensional vector space for the initial representations, where d is set to 768 in SciBERT. The matrix embedding E was processed through multiple layers of Transformer encoders, utilizing self-attention and feed-forwarded networks to generate context-aware hidden states . The hidden state of [CLS] token was chosen for the representation vector of whole citation context. For training, we used the binary cross-entropy loss , which is computed as follows:

where L and M are the number of intent category and number of sample in the dataset, and is the learnable parameter. In addition, and are the predicted and ground truth labels, if the intent jth is the intent label of the sample ith.

We trained our model on the train set of MultiCite dataset. The binary cross-entropy loss was optimized via Adam optimizer [24] with learning rate was varied in {}. The maximum number iterations for training was 5, and the batch size was set in {2,4,6,8,16,32}. The model that gives the highest performance on the development set was used for the evaluation on the test set of MultiCite dataset and generating the citation intent for our collected corpus. During the inference, to facilitate the multi-intent predictions, 0.5 was chosen as the threshold to decide whether the predicted probabilities become the predicted intent labels. We present the performance of our classifier in Section Results.

Characteristics of citation intents

In this section, we first detail the grouping of original citation intents into broader categories to facilitate field-level behavioral analysis, followed by the descriptions of the citation intent characteristics examined in our study.

Citation intent grouping

The original citation categories of the MultiCite dataset are too fine-grained. Some of intent categories are very rare and only cover minimal percentage of all citations. Fig 4 shows the distribution of all citation intents across our entire corpus.

thumbnail
Fig 4. Distribution of all citation intents across the entire corpus.

https://doi.org/10.1371/journal.pone.0358327.g004

As shown in Fig 4, there are four citation intents account for very small proportions (below 2%) are Differences, Motivation, Future work, and Extends. Therefore, it is difficult if we use them to gain insights or draw the findings about the characteristic of citation intents in the field level. To draw meaningful discipline-level insights, and motivated from prior study on citation intent grouping, we consolidated above seven citation intents into four discrete citation intent groups based on their core rhetorical objectives such as:

  • Background and Motivation [3]: the foundation-building category, comprising Background and Motivation intents to establish the research context
  • Uses and Extends [25,26]: methodological application category, representing the combined frequency of Uses and Extends intents, which reflects methodological inheritance and the practical adoption of existing techniques
  • Similarities and Differences [2]: the comparative evaluation category, which merges Similarities and Differences to evaluate the alignment or divergence of results
  • Future work: the future-oriented category, focused on the Future work intent.

Inter-category distribution of citation intents

Citation intent reflects the rationale behind why a scientific paper is referenced, in other words, it represents the mechanism through which the knowledge of one study is recognized by another. At a disciplinary scale, the high prevalence of certain citation intents reflects the priorities of that field, as well as the systematic ways in which knowledge within that domain is acknowledged and used. We measure the inter-category distribution of citation intent and perform a cross-disciplinary comparison. In this context, the terminology ’inter-category distribution’ refers the proportion of a specific intent relative to the total number of citation intents within a given field.

Average number of citation intents per article

Different to the inter-category distribution of citation intent in above, the average number of citation intents per cited paper for each field reflects the actual frequency of usage and, to some extent, the unique citation culture inherent to specific disciplinary. A similar empirical approach was adopted by Zhang et al. [27], which utilized the ‘Average cited intensity’ metric to facilitate cross-disciplinary comparisons. First, the average number of citation intent jth for the article ith is calculated as follows:

(1)

where denotes the total number of citing papers of the paper ith, represents the number of citations that the kth citing paper makes to the cited paper with citation intent , denotes the average number of citations that the cited paper receives with citation intent per citing paper. Finally, the average number of citation intents per article is calculated as follows:

(2)

where N is the total number of cited papers in the field being analyzed.

Age distribution of citation intents

The age of citation is defined as the temporal span between the publication year of the citing article and that of the cited article. Building upon this definition, prior studies indicated that there are significant differences across multiple research fields in both the mean citation age [28] and the age distribution of citation [29]. Motivated from that, in this study, we introduce the age distribution of citation intent, which captures the probability distribution of ages for specific intents. Unlike the general citation age distribution, by using the citation intent, our metric provides a granular views of how researchers’ engagement with a paper’s knowledge evolves over time. Fig 5 and 6 show the citation age distribution within first 20 years for two disciplines Agriculture and Area for the group of Background and Motivation, group of Uses and Extends, and the group of Differences and Similarities, respectively. Each plotted value represents the mean across 1,000 sampling iterations (n = 50 per iteration). A first observation is that significant disparities may exist in the age distributions among different intent categories within the same field, which indicates the need to analyze citation age through the lens of specific rhetorical functions, an insight that traditional measures fail to provide. In addition, initial observations suggest potential variations between Agriculture and Area across these intent groups, highlighting the need to explore whether such patterns differ across research fields.

thumbnail
Fig 5. Citation age distribution of three groups: Background and Motivation, Uses and Extends, and Differences and Similarities in the field Agriculture.

https://doi.org/10.1371/journal.pone.0358327.g005

thumbnail
Fig 6. Citation age distribution of three groups: Background and Motivation, Uses and Extends, and Differences and Similarities in the field Area.

https://doi.org/10.1371/journal.pone.0358327.g006

Similarity of the age distributions of citation intents

As mentioned above, distinct variations exist among the age distributions of different citation intent categories. However, the extent of these disparities fluctuates across disciplines; specifically, the distributional similarity observed in the field Agriculture is significantly lower than that in the field Area. These preliminary observations motivate a detailed comparative analysis across disciplines regarding the similarity of age distributions for citation intent. To quantify the similarity of the age distributions of citation intent, we employed Cosine Similarity and Jensen-Shannon Divergence (JSD) [30]. For each scientific field, we calculated the similarity for all possible pairs of intent group. Subsequently, these values are averaged to derive a single mean similarity score. We call this metric is inter-intent similarity. Furthermore, for each citation intent group, we measure its similarity to the overall citation age distribution within each discipline. This metric can be regarded as a normalization of the citation age distribution, as it mitigates the disciplinary differences in citation age distribution across fields. Therefore, it can reveals disparities that stem fundamentally from in the citation intents. We call this metric is intent-to-global similarity.

Predictability of the number of citation intents

Prior research has demonstrated that the number of citation serve as a robust indicator for predicting their own future trajectories [3133]. In this study, we extend these compelling findings to the context of citation intent at a cross-disciplinary scale. Specifically, we quantify the predictability of the number of citation for each intent group and perform a comparative analysis across various academic fields. In previous studies, the predictability of citation counts was assessed by calculating the correlation coefficient between two periods, t1 and t2 (where ), reflecting the predictive power of number of citation from an early stage to a subsequent point in time. Building upon this methodology, we varied t1 at several initial intervals ( years), and t2 ranging from t1 + 1–20 years. For each fixed t1, the line connecting the correlation coefficients across the values of t2 illustrates the fluctuation in predictability as a function of the prediction time (t2). The fields that yield higher curve indicates higher predictability of the number of citation intent. To ensure statistical stability, this study only calculates the correlation coefficients for fields with more than 10 documents that received at least one citation during the t1 period.

Results

Citation intent classification performance

Table 2 illustrates the classification performance of our model compared to the baseline presented in Lauscher et al. [4]. To evaluate the performance, we employed both a strict version, which requires an exact match between the entire predicted set and the gold standard, and a weak version, where a prediction is considered successful if there is at least one overlapping label. The results on the development set demonstrated high accuracy, yielding scores of 0.62 and 0.77 for the strict and weak versions, respectively. Furthermore, our model achieved highly competitive performance on the test set, reaching a strict accuracy of 0.66 and a weak accuracy of 0.79, which is competitive with that of the baseline’s (0.78).

thumbnail
Table 2. The comparison of classification performance on the MultiCite dataset.

https://doi.org/10.1371/journal.pone.0358327.t002

Table 3 reveals the classification performance on MultiCite dataset and the number of training samples for each citation intent. Overall, the results demonstrate a consistent performance across the all intent categories on the MultiCite dataset, with F1-scores ranging from 0.5133 to 0.8321. The proposed model achieves peak performance on majority classes, including Background, Uses, and Differences with 0.83, 0.77, and 0.70 F1-score, respectively. Despite the relative scarcity of training data for intents such as Extends, Motivation, Similarities, and Future work, the model still maintains good performance on these classes. The Future work category achieves an F1-score of 0.52, a acceptable result given its extremely limited sample size.

thumbnail
Table 3. The classification performance on the MultiCite test set and total of training instance for each citation intent.

https://doi.org/10.1371/journal.pone.0358327.t003

The overall test set results and the granular performance for each citation intent on the MultiCite test set confirm that our model achieves competitive results relative to existing baselines and is well-trained across all citation intent categories. Therefore, we used this model for the predictive and analytical phases of this study.

Error analysis on citation intent classification

In this section, we analyze the error pattern of our classification model, the impact of classification error across different disciplines and on our study analyses.

We first computed the confusion matrix on MultiCite test set, normalized the values into percentages to represent the error rate for each intent category and present them in Table 4.

thumbnail
Table 4. Normalized confusion matrix of our classifier on MultiCite test set.

https://doi.org/10.1371/journal.pone.0358327.t004

Based on this matrix, we can see that the intent categories most frequently misclassified are Future work (52.4%) and Extends (46.1%). In additional, the most frequent misclassification patterns identified are Extends to Uses (25.1%), Future work to Background (23.8%), Similarities to Uses (16.9%), Similarities to Differences (14.88%), Future work to Motivation (14.29%).

Regarding the impact of classification patterns across disciplines, we conducted a qualitative examination of incorrectly assigned labels on five above misclassification pairs, and showed them in Table 5.

thumbnail
Table 5. Qualitative examples of misclassifications on the MultiCite dataset.

https://doi.org/10.1371/journal.pone.0358327.t005

From the table, we can observe that the classification errors primarily arise from semantic ambiguity inherent in scientific writing in general. Particularly, in the first example, the model confuses the utilization of a baseline system with extending it using a deterministic approach (Extends to Uses). In the second sample, the classifier confuses describing a solution as merely providing background context (Background) with identifying it as a potential future direction (Future work). In the third example, describing a component derived from prior work causes the model to confuse methodological similarity with the utilization of an existing resource (Similarities to Uses). In the fourth sample, the model fails to identify the primary comparative intent of the sentence, confusing similarity indicators like ’comparable’ with contrastive elements such as ’better’ (Similarities to Differences). In the final sample, the classification model confuses establishing research motivation through the high performance of a prior model (Motivation) with intending to build upon it as a direction for future work (Future work). Because these errors primarily stem from semantic ambiguity inherent in general academic citation style rather than domain-specific factors, these misclassifications patterns are expected to remain largely consistent when applying the model across different research disciplines.

Regarding the influence of such uncertainty on the robustness of our research’s conclusion, we analyze the impact of each above error type on the analytical results across the field as follows:

  • Extends to Uses: In our analysis, we group Uses and Extends into a single category due to their shared rhetorical functions; thus, misclassifications between them do not significantly impact our field-level findings.
  • Future work to Background and Future work to Motivation: Because the intent Future work constitutes a very small proportion of our corpus (only 0.29%, see Fig 4), the impact of any misclassifications associated with it on our overall analysis remains negligible. Additionally, we conducted a human evaluation experiment (see Table 6 and Subsection Human evaluation) to assess the alignment between our classifier’s predictions and human annotations across a representative sample of our corpus. Results from Table 6 demonstrate that our classification model and human annotators achieved strong agreement () across all three labels: Future work, Motivation, and Background with agreement levels ranging from Moderate to Almost Perfect, and a large majority (50/63 samples) falling into the Substantial to Almost Perfect range. Therefore, these classification errors do not affect significantly the validity of our field-level conclusions.
  • Similarities to Differences: In our study, we grouped the Similarities and Differences intents into a single category, as both reflect comparative perspectives. Therefore, this misclassification pattern does not significantly affect our analysis.
  • Similarities to Uses: Although the misclassification rate from the Similarities to Uses reaches 17% on the MultiCite test set, MultiCite is designed as a challenging benchmark dataset for citation intent classification, and its class distribution does not reflect the citation intent distribution across broader discipline-scale. Therefore, we analyze the impact of this classification pattern according to the result from human evaluation on our corpus. Particularly, Table 6 shows that for the Similarities and Uses labels, the agreement between the classification model and human annotators consistently reaches the Substantial to Almost Perfect range. Therefore, this classification error exerts a negligible impact on our analytical findings.
thumbnail
Table 6. Cohen’s Kappa score between our classifier and human across disciplines.

https://doi.org/10.1371/journal.pone.0358327.t006

Overall, these above classification error patterns do not accounts for significant impact on our findings.

Human evaluation

In this subsection, we evaluate the predictive capability of our proposed classification model on the collected corpus. To this end, we sampled a representative dataset and engaged an academic expert to annotate it, subsequently evaluating the level of agreement between the classification model and the human annotator. Specifically, for each discipline and citation intent, we randomly selected 30 samples that were predicted by our model to contain that specific intent. To ensure high-quality annotation, the task was performed by a Ph.D. candidate in Computer Science who possesses sufficient domain knowledge to comprehend the nuanced characteristics of citation intents. The annotator was briefed on the formal guidelines and definitions of citation intents as outlined in [4]. The dataset containing both human-annotated labels and model classification outputs is made publicly available in our data repository. We use Cohen’s Kappa () [34] to measure the agreement between the model predictions and human annotations across all disciplines and citation intents, and show the results in Table 6. The interpretation of follows the work in [34].

Our results show that there is no domain and intent that gives , showing that the agreement between our classification model and humans is all at least Moderate or above. For each intent label, averaging across academic domains revealed high mean values ranging from 0.617 to 0.821, indicating that the average agreement across intents ranged from Substantial to Almost perfect. For each academic discipline, averaging across all citation intents also yielded high score, ranging from 0.703 to 0.825, indicating that the mean agreement across academic domains likewise spanned from Substantial to Almost perfect. These results demonstrate that our intent classification model performs consistently well across all citation intents and academic domains, confirming its high generalizability and robustness without domain-specific bias. Consequently, the model is fully qualified for cross-disciplinary citation analyses in this study.

Inter-category distribution of citation intents

Since the analysis of the inter-category distribution of citation intents can be sensitive to data imbalance, we adopted a resampling procedure in this analysis. Specifically, we set the sample size to Nsample = 50 (satisfying the Central Limit Theorem threshold) and performed 1,000 sampling iterations across all disciplines. Statistical testing and analysis were executed on each resampled dataset, and the results were aggregated for final reporting.

Fig 7 illustrates the mean value of inter-category distribution of citation intents across all disciplines, categorized into four groups: Background and Motivation, Uses and Extends, Differences and Similarities, and Future work. The data demonstrated that Background and Motivation is the predominant intent group across all disciplines, consistently exceeding 75% of total citations. The highest prevalence of this intent group was observed in Sociology, Philosophy, Area, and History, where they accounted for nearly 92% of all citations. In contrast, Comp Sci AI exhibited the lowest proportion for this group intent, at approximately 78%.

thumbnail
Fig 7. Inter-category distribution of four groups of citation intent across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.g007

In addition, among the remaining intent groups, the group of Uses and Extends generally exhibited a higher proportion than both group of Differences, Similarities and Future work in most instances. Within this group, fields such as Philosophy, Sociology, Psychology, Area, and History recorded the lowest values, ranging from 2% to 4%. In contrast, disciplines such as Biotechnology, and Comp Sci AI demonstrated the highest prevalence, exceeding 12%.

Following the Uses and Extends category is the group of Differences and Similarities. Within this group, fields such as Area, Philosophy, and Engineering recorded the lowest frequencies, accounting for less than 3%. Conversely, Agriculture, Economics, Linguistics, and Psychology exhibited the highest prevalence, with values exceeding 6%. Across all categories, the intent Future work consistently represented the smallest proportion of citation intents, accounting for less than 0.6% of the total distribution. Consequently, this category was excluded from further analysis to maintain focus on the dominant citation patterns between the disciplines.

We utilized a Welch’s ANOVA to verify whether any significant differences exist between these fields for each group of intent. Welch’s ANOVA was used instead of classical ANOVA to ensure robustness against potential heterogeneity of variance across fields. Table 7 reports the mean of df2, F-statistic, (), and the p-values. For p-values, we report the mean value along with the proportion of iterations with p < 0.05 across 1,000 runs.

thumbnail
Table 7. Welch’s ANOVA results for three groups of citation intent across different research fields.

https://doi.org/10.1371/journal.pone.0358327.t007

The results in Table 7 revealed that there is at least one significant differences within all examined disciplines, with p-values far below the 0.05 threshold (p < 10–19 in all cases). The number of time that p < 0.05 are 1,000 for all citation intent groups, indicates the consistent of difference through the resampling procedure. The robust F-statistics (all F > 10), underscored the substantial impact of disciplinary variation on citation behavior. Furthermore, the partial eta squared values () ranged from 0.15 to 0.16 indicate a medium effect size, confirming that research fields account for a considerable portion of the variance.

To evaluate pairwise differences across disciplines and identify clusters of highly similar fields, we conducted Games–Howell post hoc tests across all 1,000 resampling iterations alongside Welch’s ANOVA. For each intent group, we constructed two heatmaps displaying mean pairwise differences. The first heatmap marks field pairs exhibiting statistically significant differences (p < 0.05) in more than 5% of the iterations (>50 runs) with an asterisk (*), where unmarked pairs represent highly similar discipline pairs. The second heatmap marks pairs remaining significant in over 95% of the iterations (>950 runs) an asterisk (*), strictly isolating highly divergent pairs.

The post hoc analyses results of the group of Background and Motivation are shown in S1 Fig and S2 Fig, respectively. As shown in S1 Fig, there are several groups of discipline which exhibit the similarity such as (Sociology, Philosophy, Area, and History), or (Agriculture, Economics, Biotechnology). As shown in S2 Fig, the disciplinary pair exhibiting the most pronounced difference is History and Comp Sci AI, as History yields the highest proportion while Comp Sci AI records the lowest value.

The post hoc analyses results of the group of Uses and Extends are shown in S3 Fig and S4 Fig, respectively. As shown in S3 Fig, there are four disciplinary clusters demonstrated a high degree of homogeneity such as (Physics, Physical Chemistry, Medicine), (Economics, Ecology, Engineering, Film Radio TV), and the humanities-focused group (Philosophy, Psychology, Sociology, Area, History). By contrast, as shown in S4 Fig, the disciplinary pair exhibiting the largest difference is Comp Sci AI and History, as Comp Sci AI yields the highest proportion for this intent group, whereas History records the lowest.

Finally, the post hoc analyses results of the group of Differences and Similarities are shown in S5 Fig and S6 Fig, respectively. As shown in S5 Fig, a strong convergence was observed within two specific sets: (Mathematics, Multidisciplinary, Plant) and (Area, History) while S6 Fig suggests that Agriculture and Philosophy exhibit the greatest divergence, recording the highest and lowest values, respectively.

Average number of citation intents per article

Similar to the inter-category analysis, we applied the resampling procedure to evaluate intra-category variations. Fig 8 shows the mean of the average number of citation intents per article for 21 disciplines in three intent groups across all iterations.

thumbnail
Fig 8. Average number of citations by three citation intent groups across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.g008

The results indicated that among the three citation intent categories, Background and Motivation consistently yielded the highest values across most cases, with an average of 1–2 citations per referenced document. Following it, the Uses and Extends group exhibits a significantly lower frequency, averaging between 0.04 and 0.3 citations per document. Similarly, the group of Differences and Similarities recorded the lower number than that of group of Background and Motivation, with a mean from 0.05 to 0.18 citations per cited article. This results align with previous observations in the inter-category intent distribution, where the group of Background and Motivation persistently accounts for a dominant proportion. Besides, intuitively, this metric estimates the average number of citations per intent category on per citing paper, that is, when a paper is cited by specific citing paper, how many times a specific citation intent appears on this citing paper. Therefore, our results with the values of 0–2 in three citation intent groups are reasonable, consistent with the characteristic of our dataset, as a citing paper usually cites a target paper between 0 and 2 times for any given intent category.

To examine whether significant differences exist within each citation intent category across various disciplines, a Welch’s ANOVA was conducted. The statistical outcomes of this analysis are summarized in Table 8 with similar manner to inter-category analysis.

thumbnail
Table 8. Welch’s ANOVA results for average number of citation intents per document.

https://doi.org/10.1371/journal.pone.0358327.t008

The analysis revealed at least one statistically significant difference among the examined categories, with p-values far below the 0.05 threshold (p < 10–15 in all cases). The number of time that p < 0.05 are 1,000 for all citation intent groups, indicate the consistent results through the resampling procedure. Besides, the magnitude of these differences, as measured by the partial eta squared () also are significant. The most pronounced interdisciplinary variation is observed in the group of Background and Motivation (), suggesting that there are substantial disciplinary variations in the use of citations for providing research context. Specifically, Linguistics (1.96), Film Radio TV (1.96), and History (1.90) exhibit the highest average citation counts, which are substantially greater than those in fields such as Comp Sci AI (1.25) and Physics (1.34). The maximum gap between these disciplines reached a substantial margin of approximately 0.71 citations per document.

A similar trend in significant differences across fields was observed in the group of Uses and Extends but with a slightly lower effect size of . Particularly, Comp Sci AI (0.30), Biotechnology (0.21), and Mathematics (0.21) emerged as the leading fields, significantly outperforming Philosophy (0.06), Psychology (0.07), Sociology (0.06), Area (0.06), and history (0.04). The maximum discrepancy of approximately 0.25 indicates that the propensity to extend prior research is nearly seven times higher in Comp Sci AI than in History, highlighting a major divergence in disciplinary methodologies.

In addition, the group of Differences and Similarities intent exhibited a degree of interdisciplinary variation (), slightly higher than that of group of Uses and Extends. The primary discrepancies within this category arise between the highest-ranking disciplines such as Agriculture (0.18) with the lowest engagement, including Philosophy and Engineering (0.05), showing maximum of 0.13 citations per document.

Additionally, we conducted Games–Howell post hoc tests to identify similar and distinct disciplinary pairs, following the same procedure as the inter-category analysis. The post hoc analyses results of the group of Background and Motivation, group of Uses and Extends, and the group of Differences and Similarities are shown in (S7 Fig, S8 Fig), (S9 Fig, S10 Fig), and (S11 Fig, S12 Fig), respectively. In the group of Background and Motivation, S7 Fig shows several discipline clusters demonstrate high structural proximity are (Linguistics, Film Radio TV, History), (Biotechnology, Physics), and (Area, Sociology) while S8 Fig shows that three disciplines such as Linguistics, Film Radio TV, and History exhibit the most pronounced differences compared to Comp Sci AI, as Comp Sci AI records the lowest value for this measure.

In the group of Uses and Extends, while S9 Fig shows several high proximity discipline clusters such as (Philosophy, Psychology, Sociology, Area), (Comp Sci AI, Biotechnology, Mathematics), and (Economics, Ecology, Engineering), S10 Fig indicates that Comp Sci AI and History is the most divergent pair, as Comp Sci AI records the highest value while History yields the lowest.

In the group of Differences and Similarities, while S11 Fig shows several high structural proximity cluster as (Philosophy, Engineering) and (Materials, Mathematics, Medicine), S12 Fig indicates that Agriculture exhibits the greatest divergence when compared with Engineering and Philosophy, as Agriculture records the highest value while the latter two yield the lowest.

Age distribution of citation intents

For this experiment, we applied the same resampling framework with 1,000 iterations and a sample size of 50 per iteration. Fig 9 presents the average of the citation age distribution across 20 post-publication years for the three intent groups, calculated over all resampled iterations. A consistent trend emerged across the three charts: citation counts steadily climb to a peak between 2 and 10 years before entering a gradual decline. Consequently, the vast majority of citations was concentrated within the first 15 years of a paper’s lifecycle. Notably, after 15 years, citation probabilities across all intent groups remained consistently above zero. This ’long tail’ indicates that scientific contributions retain their relevance and continue to be revisited over the long term. This pattern demonstrates a consistency between the distribution of individual intent groups and the overall citation age patterns established in prior literature [35].

thumbnail
Fig 9. The citation age distribution of three groups of intent: Background and Motivation, Uses and Extends, and Differences and Similarities over 20-year period for 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.g009

Despite these shared characteristics, a granular analysis of interdisciplinary variance revealed distinct behavioral nuances across intent categories. The Background and Motivation group exhibited the highest level of cohesion, with the trajectories of all 21 fields closely aligning with minimal divergence compared to two remaining intent groups. In contrast, inter-field disparities became more pronounced within Differences and Similarities, as well as Uses and Extends, underscoring that these intent groups are more domain-dependent.

To rigorously validate these disparities on each citation intent group, we conducted Chi-square () tests of independence to confirm the statistical significance of the distributional variations observed across different intent groups and scientific fields. We present the Cram’er’s V coefficients and statistical significance p-values for all intent groups. The interpretation of the strength of association follows the criteria established by Lee [36].

Table 9 shows the mean of Cramér’s V coefficients and statistical significance p-value result over all iterations for the group of Background and Motivation. If the proportion of iterations yielding p < 0.05 exceeds 50%, the pair of disciplines exhibits a statistically significant difference. Conversely, a proportion below 50% indicates no statistically significant difference between the pair. This way of evaluation is applied for all citation intent groups. The results from Table 9 shows that almost pairwise comparisons yielded p < 0.05. This indicates that the citation patterns of the group of Background and Motivation are significantly different and dependent on the specific discipline. In addition, the magnitude of these differences varied across disciplinary pairs. Specifically, 4.29% (9 pairs) fell within the Negligible range, followed by 54.29% (114 pairs) classified as Weak, 33.81% (71 pairs) as Moderate, 7.14% (15 pairs) exhibited a Relatively Strong, and 0.48% (1 pairs) for Very Strong. The disciplinary pair exhibiting the largest difference for this intent is Physical Chemistry and Psychology with the Cramér’s V is 0.61 (Very Strong).

thumbnail
Table 9. Pairwise Cramér’s V coefficients and statistical significance of age distribution of group Background and Motivation across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.t009

A large number of intent pairs falling within the Negligible range is consistent with the distributions illustrated in Fig 9, where the majority of fields exhibited similar shapes. The only strong divergence within this intent group arises from Psychology, which exhibited several Cramér’s V values between 0.38 and 0.61 when compared to other disciplines, reflecting a unique citation tendency in this field compared to other scientific disciplines. This trend was also visually corroborated in Fig 9, which showed that Psychology exhibits a longer citation peak latency compared to other fields, with the peak occurring between 15 and 17 years.

Table 10 shows the mean of Cramér’s V coefficients and statistical significance p-value result for the group of Uses and Extends over all iterations. Similar to Background and Motivation group, this group also exhibits existence of statistical divergence across disciplinary pairs, with 179 over 210 pairs yield p < 0.05, account for 85.24%. In addition, the magnitude of the differences based on Cram’er’s V coefficients. Particularly, consider only on 179 different pairs, there are only 1 pairs (0.56%), 32 pairs (17.88%) in Weak, 135 pairs (75.42%) in Moderate, and 11 pairs (6.15%) in Relative Strong. Within the different discipline pairs, the average Cramér’s V score of this intent group is 0.259, higher than that of the group of Background and Motivation (0.206). This value suggested that the degree of disciplinary variation for the Uses and Extends intent group is higher than that of Background and Motivation group. For this intent, the disciplinary pair exhibiting the largest difference is Physical Geography and Psychology, with a value of Cramér’s V is 0.49 (Relative Strong).

thumbnail
Table 10. Pairwise Cramér’s V coefficients and statistical significance of age distribution of group Uses and Extends across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.t010

Table 11 reveals the mean of Cramér’s V coefficients and statistical significance p-value result for the group of Differences and Similarities over all iterations. Similar to two above intent groups, the group of Differences and Similarities also exhibits the existence of statistical difference across most pairs of scientific field with 80% of different discipline pairs (168 over 210). Besides, the Cram’er’s V coefficients are vary with 44 pairs (26.19%) was classified as Weak, 111 pairs (66.7%) in Moderate, and 13 pairs (6.2%) in Relative Strong. Within the different discipline pairs, the average Cramér’s V score is 0.247, which slightly lower than that of Uses and Similarities and greater than that of group of Background and Motivation, suggested that the degree of disciplinary variation of this intent group is at a moderate level across three intent groups, which consistent with the observation in Fig 9. For this intent group, the disciplinary pair exhibiting the largest difference is Biotechnology and Psychology, with a value of Cramér’s V is 0.53 (Relative Strong).

thumbnail
Table 11. Pairwise Cramér’s V coefficients and statistical significance of age distribution of group Differences and Similarities across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.t011

Similarity of the age distribution of citation intent

Inter-intent similarity

We presents a comparative analysis of citation age distributions of intent across 21 scientific disciplines, quantified via JSD and Cosine similarity in Fig 10. The reported values represent the mean of the measurements obtained across 1,000 iterations. Overall, the results from Fig 10 reveals a degree of similarity across 21 disciplines where most Cosine similarity score remained predominantly above 0.89 and most JSD values consistently fell below the 0.07. Such results underscored a remarkable consistency in three group of citation intent practices within respective research domains, suggesting that the age distribution of citation intent are fundamentally stable within most scholarly ecosystem.

thumbnail
Fig 10. Comparative analysis of JSD and Cosine similarity across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.g010

Despite a consistently high Cosine similarity and low JSD were shared across most academic domains, a pronounced inter-disciplinary variance is still observed. Particularly, fields such as Film Radio and Multidisciplinary exhibits the highest levels of internal consistency, with JSD values lower than 0.01 and Cosine similarity greater than 0.96. Conversely, a distinct divergence is observed in Area and History, where JSD spikes above 0.044 and Cosine similarity diminishes below 0.9, where Area emerges as the most prominent outlier (JSD = 0.064, Cosine = 0.893).

To examine whether significant differences in this characteristic across disciplines, a Welch’s ANOVA was conducted. The statistical outcomes of this analysis are summarized in Table 12.

thumbnail
Table 12. Welch’s ANOVA results for inter-intent similarity in the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.t012

The results from Table 12 reveals the existence of differences for both metrics, with p-values far below the 0.05. Besides, the magnitude of these differences, as measured by the partial eta squared () also are significant, ranging from 0.367 to 0.643. Additionally, two post hoc tests were also conducted to examine specific pairwise differences across disciplines. The post hoc test results for JSD and Cosine similarity are presented in S1 Table and S2 Table, respectively. As shown in S1 Table and S2 Table, largest differences originate from Area Studies and History relative to the remaining disciplines, which aligns with the observations in Fig 10.

In addition, methodologically, the parallel deployment of these metrics yields highly consistent results, evidenced by a near-perfect negative monotonic correlation (Spearman’s ). This robust alignment underscores the validity and methodological consistency of the analytical framework employed in this study.

Intent-to-global similarity

Fig 11 shows the mean of JSD between the citation age distribution between specific intent group with each particular disciplines across all sampling iterations. The results indicated that among the three citation intent categories, Background and Motivation consistently yielded the lowest values across most cases, ranging from 0 to 10–3. The group of Uses and Extends and the group of Differences and Similarities recorded the higher JSD, ranging from 0.02 to 0.0.063 and 0.003 to 0.04.

thumbnail
Fig 11. Similarity between each citation intent groups with overall citation age distribution across 21 scientific fields.

https://doi.org/10.1371/journal.pone.0358327.g011

To examine whether significant differences exist within each citation intent group across disciplines, a Welch’s ANOVA was conducted. The statistical outcomes of this analysis are summarized in Table 13.

thumbnail
Table 13. Welch’s ANOVA results for intent-to-global similarity in the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.t013

The analysis reveals differences for all examined categories, with p-values far below the 0.05. Besides, the magnitude of these differences, as measured by the partial eta squared () also are significant, ranging from 0.395 to 0.7863, suggesting that there are disciplinary variations in all citation intent groups. Specifically, in the group of Background and Motivation, Biotechnology gives the highest JSD (0.001) while the Engineering gives the losest JSD (only 0.00002). In the group of Uses and Extends, Area gives the highest JSD while Mathematic gives the lowest. In the group of Differences and Similarities, Area still is the filed has the highest JSD while Plant gives the lowest JSD.

We also conducted and showed the results of a Games-Howell post hoc analysis to verify pairwise differences across disciplines. The post hoc analyses results for group of Background and Motivation, Uses and Extends, Differences and Similarities are shown in S3 Table, S4 Table, and S5 Table, respectively. For the Background and Motivation group (S3 Table), the most pronounced differences stemmed from Biotechnology, followed by Plant Sciences, relative to other fields. In contrast, both the Uses and Extends (S4 Table) and Differences and Similarities (S5 Table) groups exhibited a shared pattern, with Area Studies driving the largest cross-disciplinary discrepancies, followed by History.

Predictability of the number of citation intents

Fig 12 illustrates the predictive power across 21 disciplines for all citation intent groups, averaged across all sampling runs. A consistent downward trend in correlation levels was observed across all groups when increasing t2 while keeping t1 fixed. This decay was most pronounced in the immediate aftermath of the prediction point, followed by eventual stabilization as t2 extended further. Furthermore, increasing the value of t1 while holding t2 constant led to higher correlation coefficients, confirming that a larger historical data foundation provides enhanced informational cues for predicting future scientific achievement. It is also noteworthy that while correlation coefficients vary significantly across disciplines, these gaps are most visible when t1 is small and tend to diminish as t2 increases.

thumbnail
Fig 12. Predictability of the numbef of citation intents across 21 fields over three prediction intervals for three citation intent groups.

https://doi.org/10.1371/journal.pone.0358327.g012

Regarding the Background and Motivation intent group (Fig 12A), all research fields demonstrated remarkably high initial predictive power during the first five-year interval (r > 0.80). However, while every discipline followed a consistent downward trend toward the 19th year, the magnitude of this decay varied significantly across fields, spanning a wide range from 0.87 to 0.26. The predictive power of disciplines such as Sociology, and History reached its nadir 19 years after publication. Sociology, in particular, witnessed a dramatic collapse in correlation, with the correlation value r plummeting by 0.69, descending from an initial 0.95 to a mere 0.26, with averaged r computed for whole time is 0.537. In contrast, fields such as Biotechnology, Multidisciplinary, Plant, Medicine, and Physics maintained the high predictive power, with r remaining above 0.80 at final time. More specificially, Biotechnology emerged as the most stable discipline, exhibiting a marginal decline of only 0.12, dropping from 0.99 to 0.87, and averaged r for whole time is 0.92.

Extending the t1 interval from 5 to 7 years yielded a general improvement in predictive power across all disciplines. Furthermore, this shift led to a noticeable convergence between fields, as the disparity narrowed to a range between a minimum of 0.37 and a maximum of 0.93. Biotechnology, Medicine, Plant, and Physics maintained their high predictability with r-values consistently exceeding 0.90. These fields exhibited remarkable stability, with their predictive power declining by less than 0.10 from the baseline. Conversely, the lowest correlations were found in several disciplines such as Philosophy, Sociology, and History, where predictive values failed to reach the 0.50 threshold at the final time.

Extending the t1 window to the first 10 years after publication significantly bolstered predictive power across all disciplines, a result of the expanded observational data. By the final time frame, all fields exceeded a correlation threshold of 0.60. Biotechnology, Medicine, and Plant have the r-values surpassing 0.93, whereas Philosophy, Psychology, and Sociology represented the lower tier, with results ranging between 0.62 and 0.71.

Table 14 presents the Welch’s ANOVA test results for the mean r values calculated across the entire t2 time span for all three stages (t1 = 3, 5, 7). Note that to ensure statistical robustness, correlation coefficients were computed exclusively for fields containing more than 10 cited documents and over 20 cumulative citations during t1 during each sampling iteration.

thumbnail
Table 14. Welch’s ANOVA results for averaged r for whole time t2 with different t1 for group of Background and Motivation.

https://doi.org/10.1371/journal.pone.0358327.t014

Consistent with our observations above, Table 14 shows the p-value was consistently below 0.05 and the () exhibited a large effect size across all cases, indicates the existence of significant variations across disciplines. Additionally, post hoc tests were conducted for t1 = 5, 7, 10, with the corresponding results presented in S6 Table, S7 Table, and S8 Table, respectively. Specifically, at t1 = 5 S6 Table, the most substantial mean difference is observed between Plant and Sociology (). As the observation window expands to t1 = 7 (S7 Table), the maximum divergence shifts to the pair of Biotechnology and History (). Finally, for t1 = 10 (S8 Table), the greatest disparity is identified between Biotechnology and Psychology ().

The correlation results for the citation intent group of Uses and Extends are presented in Fig 12B. During the initial five-year interval, similar to the trends observed in the Background and Motivation intent groups, various research fields exhibited a high degree of divergence, ranging significantly from 0.42 to 0.91. Philosophy emerged as the discipline with the weakest predictive performance. Notably, by the 19th year, the correlation coefficient for Philosophy reached a mere 0.42, representing a sharp decline of 0.42 compared to its initial values at the onset of the t2 period, and only obtains 0.533 for average in whole time. In contrast, Plant emerged as the top field, yielding correlation coefficients of 0.91, with averaged r as 0.95 in whole time. Expanding the t1 window to 7 years post-publication resulted in a universal increase in correlation coefficients across all disciplines, accompanied by a noticeable trend toward convergence. Under this setting, Philosophy remained the notable underperformers, with r-values of 0.54 at the final time and averaged r as 0.651 in whole time. Conversely, Biotechnology is the leading field, achieving a peak correlation of 0.96. Under the 10-year t1 configuration, the correlation coefficients across all disciplines exhibited a high degree of convergence, with minimal interdisciplinary variance. Despite this overall uniformity, Material, Philosophy and History continued to yield the low results, recording r-values lower than 0.8 at the final time. Conversely, the peak predictive performance was observed in several disciplines such as Biotechnology, Plant, and Area, all of which achieved a near-perfect correlation of 0.98.

Table 15 presents the Welch’s ANOVA test results for the mean r values calculated across the entire t2 time span for all t1 = (5, 7, 10).

thumbnail
Table 15. Welch’s ANOVA results for averaged r for whole time t2 with different t1 for group of Uses and Extends.

https://doi.org/10.1371/journal.pone.0358327.t015

As shown in Table 15, the p-value was consistently below 0.05 and the () exhibited a large effect size across all cases, indicates the existence of significant differences across disciplines, which is consistent with our observations above. Additionally, post hoc tests were conducted for t1 = 5, 7, 10, with the corresponding results presented in S9 Table, S10 Table, and S11 Table, respectively. Specifically, at t1 = 5 (S9 Table), the most substantial mean difference is observed between Biotechnology and Philosophy (). As the observation window expands to t1 = 7 (S10 Table), the maximum divergence still is witnessed between Biotechnology and Philosophy (). Finally, for t1 = 10 (S11 Table), the greatest disparity is identified between Biotechnology and Materials ().

The predictive power of all disciplines for the group of Differences and Similarities was shown in Fig 12C. In the 5-year t1, similar to the two above citation intent groups, a high divergence was witnessed across all scientific fields. The discipline exhibited significantly weaker predictive performance compared to their counterparts is Sociology, which recorded the lowest correlation values of 0.43 at the final time, and averaged r is only 0.55. By contrast, Biotechnology is the best field of predictive accuracy across all disciplines as it recorded a value of 0.91 after 19 years and averaged r is 0.94. As t1 was extended to 7 and 10 years post-publication, the predictive power across all disciplines generally improved, and the performance gap between them tended to narrow. In both the t1 = 7 and t1 = 10 settings, no single discipline exhibited a significantly lower performance compared to the others, while Biotechnology and Plant lead. These fields obtain averaged correlation coefficients in whole time ranged from 0.96, 0.95 and 0.99, 0.97, respectively.

Table 16 presents the Welch’s ANOVA test results for the mean r values calculated across the entire t2 time span for all t1 = (5, 7, 10).

thumbnail
Table 16. Welch’s ANOVA results for averaged r for whole time t2 with different t1 for group of Differences and Similarities.

https://doi.org/10.1371/journal.pone.0358327.t016

As shown in Table 16, the p-value was consistently below 0.05 and the () exhibited a large effect size across all cases, indicates the existence of significant differences across disciplines, and consistent with above observations. Besides, post hoc tests were conducted for t1 = 5, 7, 10, with the corresponding results presented in S12 Table, S13 Table, and S14 Table, respectively. Specifically, at t1 = 5 (S12 Table), the most substantial mean difference is observed between Biotechnology and Sociology (). As the observation window expands to t1 = 7 (S13 Table), the maximum divergence still is witnessed between Biotechnology and History (). Finally, for t1 = 10 (S14 Table), the greatest disparity is identified between Biotechnology and History ().

Discussion

Differences in the characteristics of citation intents across scientific disciplines

In this study, we investigate the characteristic of citation intent in discipline level by examining five measurements. Based on the experimental results, we found out that while significant variations persist in each characteristic of citation intent group across the scientific disciplines, certain fields simultaneously exhibit notable similarities in their characteristics.

The inter-category distribution of four groups of citation intent suggests the existence of difference in how the researcher in specific discipline recognize the prior research. Philosophy, Sociology, Area, and History exhibited the highest proportion of citations within the Background and Motivation categories. This suggests that scholarly articles in these fields require more extensive space to establish research contexts and rationales. Furthermore, these above disciplines and Psychology demonstrated the lowest prevalence of Uses and Extends intents. This finding indicates that the inheritance and reuse of existing methodologies are limited within these domains. Interestingly, our results are consistent with the hypotheses and findings presented in [37,38], where the author stated that due to low consensus, social science researchers often must spend more space to providing research context and justifying their rationale and approach. The low ratio of reuse and inheritance of methodologies here can also be considered a consequence of the lack of consensus. However, while their study relied on relatively simple metrics such as article length or the number of references, we directly exploit the rhetorical functions of citations through citation intent. A notable exception was observed in the fields of Comp Sci AI, which exhibited the lowest frequency of Background and Motivation intents while utilizing a high proportion of Uses and Extends citations. These result accurately reflect the characteristics of Com Sci AI as a field primarily focus on the construction, refinement of academic methods and exhibits a high degree of scholarly inheritance. This inherently increases the proportion of citation intents like Uses and Extends while reducing the need for citations dedicated to establishing the research context like Background or Motivation.

The results on average number of citation intent per article also reveal disciplinary disparities. Similar to the inter-category distribution, we observe a contrast in the results for Comp Sci Ai and several disciplines such as Philosophy, Sociology, Area, and History. This further confirms that different scientific fields exhibit differences in how they inherit methodologies from predecessor studies. Specifically, Comp Sci AI demonstrates a strong methodological inheritance, whereas fields such as Philosophy, Sociology, Area, and History show very limited consensus.

The age distribution of citation intent groups also reveals variations across different fields of science in all intent group, demonstrating that several distinct fields possess unique pattern of knowledge consumption over time. The disciplinary variations in the age of citations across multiple disciplines were explored in [28,29]. Although Wahle et al. [28] measured the average citation age, we believe that differences in the mean of citation age is equal to the differences in age distribution of citation. Our study extends this idea from raw citation to citation intents and also find out similar findings. In addition, we also find out another interesting outcome that while all three intent groups exhibit variations, the magnitude of these disparities differs across intent group. Specifically, the differences within the Background and Motivation groups are less pronounced compared to those observed in the remaining two intent categories. This demonstrates that the way papers reference others for technical inherit (Uses or Extends) or comparative analysis (Differences or Similarities) is a more distinctive characteristic of each specific field.

The inter-intent similarity of age distribution patterns across the three citation intent groups underscore the existence of significant disciplinary divergence and shows that Area, and History exhibit the lowest levels of similarity among the 21 scientific disciplines analyzed. Furthermore, across both intent groups Uses and Extends and Differences and Similarities, these two above disciplines exhibit low similarity to the overall citation age distribution, which is also consistent with the findings from the inter-intent similarity analysis. As the age distributions reflect how the perceived importance of a study evolves over time, these findings indicate that researchers across two above fields differ in how they recognize the contribution of prior work for specific rhetorical roles.

The analysis of the predictability of three groups of citation intent also reveals existence of disciplinary disparities across academic fields. Our results are consistent with the findings in [31], which also identified disciplinary variations in the predictability of citation counts. However, our study extends this analysis by examining these differences at a more granular level (citation intent compared to raw citation) and across a broader scope (21 fields compared to only 6 in [31]). The disparities are most pronounced in the short-term citation accumulation and gradually diminish as the accumulation period extends. Regarding the group of Background and Motivation, the lowest levels of predictability are observed in Sociology, Philosophy, and History. In addition, for the group of Uses and Extends, History emerges as the least predictable domains. Furthermore, Sociology yields the lowest predictive accuracy within the Differences and Similarities intent group. The difficulty in predicting citation intents suggests that these research fields lack consistency in how scholarly papers are evaluated, which leads to instability in how a paper’s contribution is perceived and valued by the scientific community over time. Conversely, Biotechnology and Plant exhibit the highest levels of predictability and consistency across all citation intent categories, suggesting a high degree of consensus among researchers in these fields regarding the utilization and recognition of scholarly knowledge over time.

Theoretical implications

The findings of this study offer several significant theoretical implications in the bibliometrics as follows:

  • Uncovering cross-disciplinary variations in citation intents: To the best of our knowledge, this is the first study to conduct a comparative analysis of citation intent characteristics across disciplines within the field of bibliometrics. Previous studies, whether focusing on defining taxonomies for citation intents, constructing datasets [16], or improving the quality of citation intent prediction [39,40], have only focused on understanding citation intent at an individual scale, which is only sufficient to explain why any given author cites another paper. In contrast, our study analyzes and compares citation intents at a disciplinary scale, providing crucial insights that reflect how mechanisms of knowledge flow at a significantly larger magnitude.
  • Validating the criticality of citation intent analysis: The findings of this study underscore the critical importance of investigating citation intents within the field of bibliometrics. The result obtained from Fig 79, and 12 reveal marked differences among intent groups, confirming that only examining citation characteristics is insufficient to capture the full picture of how academic knowledge is validated and utilized. Conversely, citation intents, which reflect the intrinsic nature and rhetorical functions of citations, in our view, a potent candidate for elucidating this broad picture.
  • Establishing a standardized framework for comparative analysis of citation intents: Our research provides a comprehensive and robust framework for analyzing citation intent characteristics at a disciplinary scale, offering a foundational structure that subsequent studies can readily inherit and expand upon. Specifically, our pipeline comprises four stages: Data collection, Citation intent classifier construction, Citation intent prediction and Analysis. Regarding data collection, the pivotal task involves the extraction of citation contexts. For the classifier model construction and citation intent prediction, the critical component is the development of a robust model capable of achieving high-performance predictions of citation intents. Future studies can leverage this framework to augment training datasets with larger volumes of data or a broader range of disciplines. Furthermore, researchers may reuse or easily refine intent classification model to achieve even higher levels of predictive quality. Additionally, our study establishes a standardized set of characteristics for analyzing and comparing citation intent properties across various scholarly entities. Although the present analysis focuses on comparing diverse disciplines, the framework is highly versatile and fully extensible to other levels of granularity, such as journals, individual authors, or academic institutions.

Practical implications

The findings of this study yield several important practical applications, as detailed below:

  • Advancing the normalization of scholarly indicators: Within the context of citation intent, we consider the goal to develop a scholarly indicators for each group of citation intent across multiple disciplines?

In scientific community, to compare the scholarly impact between multiple disciplines, Category Normalized Citation Impact (CNCI), a field-normalized metric is widely used. CNCI normalizes the number of citations of a publication to expected number of citations of the same document type, publication year, and subject category. Based on CNCI, we can calculate the scholarly impact of various academic entities such as journals, institutions, nations, or researchers, across different disciplines. Regarding the citation intent, the CNCI can be easily extended by substituting the aggregate citation count with the frequency of specific intent categories.

However, such an extension encounters a significant limitation: it fails to account for the different in the citation intent pattern from multiple disciplines. As illustrated in Fig 7, citation intent priorities vary significantly across different disciplines. For instance, when consider what is best journals in the group of citation intent Uses and Extends, journals in the discipline Comp Sci AI should be assigned a higher weight than the journals in the discipline Sociology because this intent group is more prevalent and play more important role in Comp Sci AI compared to its role in Sociology. Based on this observations, we propose a weighted CNCI with group of intent i for a document as follows:

(3)

where are the proportion of intent i in the inter-category distribution of this discipline, the specific count of intent i for the considering publication, and the mean frequency of intent i across all documents with the same document type, publication year, and subject category, respectively. Based on this , we can calculate the scholarly indicator with intent group i for the journals, institutions, nations, or researchers across different disciplines.

  • Facilitating technology trend detection and future forecasting: By analyzing the age distribution of citation intents (Fig 9), we can effectively delineate the specific developmental stages of a given technology. For each discipline, our study shows the evolution of citation intents from a paper’s publication time until it reaches its peak citation volume, a period can be understood as the its trending time with each intent types. In this time, the prevalence of Uses and Extends citation intents reflects the paper’s sustained utility and its ongoing inheritance within the scientific community. Furthermore, a high frequency of Differences and Similarities intents suggests that the publication remains a pivotal benchmark for comparative analysis. Conversely, the post-peak period suggests the obsolescence of this paper. Our comparative analysis reveals that the age distributions of citation intents exhibit distinct between scientific fields. Consequently, researchers can leverage these findings to discern whether a study remains within the current disciplinary trend and to estimate its remaining period of scholarly attention. When a paper’s age exceeds the typical peak duration of its field, it may serve as a signal for researchers to reallocate their focus toward emerging trajectories.
  • Supporting curriculum design and development: The inter-category distribution of citation intents (Fig 7) demonstrates that academic disciplines maintain distinct priorities regarding specific intent groups. Notably, a high prevalence of the Uses and Extends group within a field underscores the fundamental importance of methodologies and technical frameworks in that domain. Therefore, the curriculum in such fields should prioritize the acquisition of practical tools and methodological proficiency over the purely theoretical focus found in other disciplines.

In addition, the variations in the age distribution of citation intents across disciplines (Fig 9) indicates that each field possesses a distinct trending time for its technological trajectories. This finding underscores the need for a strategy to redesign and update curriculum for different discipline. By following the age distribution of citation intents, schools can help students keep up with new technologies and remove outdated knowledge from the curriculum.

External validity

Regarding the implications for external validity, we acknowledge that our dataset primarily characterizes highly influential publications from elite journals rather than the broader citation practices of entire disciplines. However, we argue that our study retains substantial analytical value and serves as a highly reliable reference framework for understanding citation dynamics at a discipline scale. Indeed, our primary research focus is centered on articles published in top-tier journals with high citation counts. Such publications reflect the essential, foundational knowledge that drives the development of a specific discipline. Furthermore, being featured in premier venues and accumulating substantial citations demonstrates widespread recognition and validation from the scientific community. Consequently, the citation characteristics observed in these articles accurately reflect the citation dynamics associated with the vital intellectual core, the most critical segment, of the field.

Conclusions

This study provides a first data-driven comparative analysis of citation intent characteristics across 21 distinct scientific disciplines. The analysis results over 2,345,674 citation contexts demonstrates there are significant disparities across all five examined characteristics: inter-category distribution of citation intents, average number of citation intents per article, age distribution of citation intents, similarity of the age distributions of citation intent, and predictability of the number of citation intents. Particularly, Computer Science AI prioritize methodological continuity through a high prevalence of Uses and Extends intents. In contrast, fields such as Philosophy, Sociology, and History lean heavily toward Background and Motivation, reflecting a focus on conceptual framing rather than direct methodological inheritance. Furthermore, our analysis of age distribution similarity shows that Philosophy, Psychology, Area Studies, and History exhibit the lowest levels of similarity, indicating significant disparities in how researchers within these domains validate and acknowledge academic knowledge over time. In terms of predictability, Biotechnology and Plant Science yield the highest correlation coefficients, whereas Sociology, Philosophy, and History demonstrate the lowest correlation levels.

Regarding theoretical implications, this study not only provides critical insights into cross-disciplinary variations in citation intents, thereby validating the importance of such analyses, but also establishes a standardized framework for the comparative analysis of citation intents across other scholarly entities. Beyond its theoretical contributions, the findings hold significant practical value by advancing the normalization of scholarly indicators, enhancing the precision of technology trend detection, and providing a strategic basis for discipline-specific curriculum design.

Supporting information

S1 Fig. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Inter-category distribution of citation intents for identifying highly similar disciplines.

This type of heatmap marks field pairs exhibiting statistically significant differences (p < 0.05) in more than 5% of the iterations (>50 runs) with an asterisk (*), where unmarked pairs represent highly similar discipline pairs.

https://doi.org/10.1371/journal.pone.0358327.s001

(PNG)

S2 Fig. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Inter-category distribution of citation intents for identifying highly divergent disciplines.

This type of heatmap marks pairs remaining significant in over 95% of the iterations (>950 runs) an asterisk (*), strictly isolating highly divergent pairs.

https://doi.org/10.1371/journal.pone.0358327.s002

(PNG)

S3 Fig. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of: Inter-category distribution of citation intents for identifying highly similar disciplines.

https://doi.org/10.1371/journal.pone.0358327.s003

(PNG)

S4 Fig. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of: Inter-category distribution of citation intents for identifying highly divergent disciplines.

https://doi.org/10.1371/journal.pone.0358327.s004

(PNG)

S5 Fig. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of: Inter-category distribution of citation intents for identifying highly similar disciplines.

https://doi.org/10.1371/journal.pone.0358327.s005

(CSV)

S6 Fig. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of: Inter-category distribution of citation intents for identifying highly divergent disciplines.

https://doi.org/10.1371/journal.pone.0358327.s006

(PNG)

S7 Fig. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Average number of citation intents per article for identifying highly similar disciplines.

https://doi.org/10.1371/journal.pone.0358327.s007

(PNG)

S8 Fig. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Average number of citation intents per article for identifying highly divergent disciplines.

https://doi.org/10.1371/journal.pone.0358327.s008

(PNG)

S9 Fig. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of Average number of citation intents per article for identifying highly similar disciplines.

https://doi.org/10.1371/journal.pone.0358327.s009

(PNG)

S10 Fig. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of Average number of citation intents per article for identifying highly divergent disciplines.

https://doi.org/10.1371/journal.pone.0358327.s010

(PNG)

S11 Fig. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of Average number of citation intents per article for identifying highly similar disciplines.

https://doi.org/10.1371/journal.pone.0358327.s011

(PNG)

S12 Fig. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of Average number of citation intents per article for identifying highly divergent disciplines.

https://doi.org/10.1371/journal.pone.0358327.s012

(PNG)

S1 Table. Games-Howell post hoc test result for JSD metric in the analysis of Inter-intent similarity of the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.s013

(CSV)

S2 Table. Games-Howell post hoc test result for Cosine similarity metric in the analysis of Inter-intent similarity in the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.s014

(CSV)

S3 Table. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Intent-to-global similarity in the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.s015

(CSV)

S4 Table. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of Intent-to-global similarity in the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.s016

(CSV)

S5 Table. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of Intent-to-global similarity in the age distribution of citation intent.

https://doi.org/10.1371/journal.pone.0358327.s017

(CSV)

S6 Table. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Predictability of the number of citation intents with t1 = 5.

https://doi.org/10.1371/journal.pone.0358327.s018

(CSV)

S7 Table. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Predictability of the number of citation intents with t1 = 7.

https://doi.org/10.1371/journal.pone.0358327.s019

(CSV)

S8 Table. Games-Howell post hoc test result of the group of Background and Motivation in the analysis of Predictability of the number of citation intents with t1 = 10.

https://doi.org/10.1371/journal.pone.0358327.s020

(CSV)

S9 Table. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of Predictability of the number of citation intents with t1 = 5.

https://doi.org/10.1371/journal.pone.0358327.s021

(CSV)

S10 Table. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of Predictability of the number of citation intents with t1 = 7.

https://doi.org/10.1371/journal.pone.0358327.s022

(CSV)

S11 Table. Games-Howell post hoc test result of the group of Uses and Extends in the analysis of Predictability of the number of citation intents with t1 = 10.

https://doi.org/10.1371/journal.pone.0358327.s023

(CSV)

S12 Table. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of Predictability of the number of citation intents with t1 = 5.

https://doi.org/10.1371/journal.pone.0358327.s024

(CSV)

S13 Table. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of Predictability of the number of citation intents with t1 = 7.

https://doi.org/10.1371/journal.pone.0358327.s025

(CSV)

S14 Table. Games-Howell post hoc test result of the group of Differences and Similarities in the analysis of Predictability of the number of citation intents with t1 = 10.

https://doi.org/10.1371/journal.pone.0358327.s026

(CSV)

References

  1. 1. Teufel S, Siddharthan A, Tidhar D. Automatic classification of citation function. In: Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing - EMNLP ’06, 2006. 103. https://doi.org/10.3115/1610075.1610091
  2. 2. Jurgens D, Kumar S, Hoover R, McFarland D, Jurafsky D. Measuring the Evolution of a Scientific Field through Citation Frames. TACL. 2018;6:391–406.
  3. 3. Cohan A, Ammar W, Van Zuylen M, Cady F. Structural scaffolds for citation intent classification in scientific publications. In: Proceedings of the 2019 conference of the North American chapter of the Association for Computational Linguistics: human language technologies, 2019. 3586–96.
  4. 4. Lauscher A, Ko B, Kuehl B, Johnson S, Cohan A, Jurgens D. In: Proceedings of the 2022 conference of the North American chapter of the association for computational linguistics: Human language technologies, 2022. 1875–89.
  5. 5. Abu-Jbara A, Ezra J, Radev D. Purpose and polarity of citation: Towards nlp-based bibliometrics. In: Proceedings of the 2013 conference of the North American chapter of the association for computational linguistics: Human language technologies, 2013. 596–606.
  6. 6. Pride D, Knoth P. An Authoritative Approach to Citation Classification. In: Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020, 2020. 337–40. https://doi.org/10.1145/3383583.3398617
  7. 7. Beltagy I, Lo K, Cohan A. SciBERT: A Pretrained Language Model for Scientific Text. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019. 3613–8. https://doi.org/10.18653/v1/d19-1371
  8. 8. Mercier D, Rizvi S, Rajashekar V, Dengel A, Ahmed S. ImpactCite: An XLNet-based Solution Enabling Qualitative Citation Impact Analysis Utilizing Sentiment and Intent. In: Proceedings of the 13th International Conference on Agents and Artificial Intelligence, 2021. 159–68. https://doi.org/10.5220/0010235201590168
  9. 9. Roman M, Shahid A, Khan S, Koubaa A, Yu L. Citation Intent Classification Using Word Embedding. IEEE Access. 2021;9:9982–95.
  10. 10. Qi R, Wei J, Shao Z, Li Z, Chen H, Sun Y, et al. Multi-task learning model for citation intent classification in scientific publications. Scientometrics. 2023;128(12):6335–55.
  11. 11. Ghosal T, Varanasi KK, Kordoni V. A Deep Multi-Tasking Approach Leveraging on Cited-Citing Paper Relationship For Citation Intent Classification. Scientometrics. 2023;129(2):767–83.
  12. 12. Li T, Wang J, Zhang Y, Li S, Chen L. Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning. In: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, 2025. 1541–52. https://doi.org/10.1145/3711896.3736829
  13. 13. Zhang Y, Zhao R, Wang Y, Chen H, Mahmood A, Zaib M, et al. Towards employing native information in citation function classification. Scientometrics. 2022;127(11):6557–77.
  14. 14. Berrebbi D, Huynh N, Balalau O. GraphCite: Citation Intent Classification in Scientific Publications via Graph Embeddings. In: Companion Proceedings of the Web Conference 2022, 2022. 779–83. https://doi.org/10.1145/3487553.3524657
  15. 15. Du X, Ahrabian K, Ananthan ABS, Myloth RD, Pujara J. Citation intent classification through weakly supervised knowledge graphs. In: Proceedings of the Workshop on Scientific Document Understanding (SDU@AAAI), 2023.
  16. 16. Xu X, Xie Y, Zhao X, Liu Y. Mf-cite: citation intent classification in scientific papers based on multi-feature fusion. Scientometrics. 2025;130(8):4465–93.
  17. 17. Lahiri A, Sanyal DK, Mukherjee I. CitePrompt: Using Prompts to Identify Citation Intent in Scientific Papers. In: 2023 ACM/IEEE Joint Conference on Digital Libraries (JCDL), 2023. 51–5. https://doi.org/10.1109/jcdl57899.2023.00017
  18. 18. Shi S, Hu K, Xie J, Guo Y, Wu H. Robust scientific text classification using prompt tuning based on data augmentation with L2 regularization. Information Processing & Management. 2024;61(1):103531.
  19. 19. Koloveas P, Chatzopoulos S, Vergoulis T, Tryfonopoulos C. Can LLMs Predict Citation Intent? An Experimental Analysis of In-Context Learning and Fine-Tuning on Open LLMs. In: Proceedings of the 29th International Conference on Theory and Practice of Digital Libraries (TPDL), 2025. 207–24. https://doi.org/10.1007/978-3-032-05409-8_13
  20. 20. Birkle C, Pendlebury DA, Schnell J, Adams J. Web of Science as a data source for research on scientific and scholarly activity. Quantitative Science Studies. 2020;1(1):363–76.
  21. 21. Rose ME, Kitchin JR. pybliometrics: Scriptable bibliometrics using a Python interface to Scopus. SoftwareX. 2019;10:100263.
  22. 22. Priem J, Piwowar H, Orr R. OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. https://openalex.org 2022.
  23. 23. Kinney R, Anastasiades C, Authur R, Beltagy I, Bragg J, Buraczynski A. The semantic scholar open data platform. https://www.semanticscholar.org 2023.
  24. 24. Kingma DP, Ba J. Adam: A Method for Stochastic Optimization. In: Proceedings of the 3rd International Conference on Learning Representations (ICLR), 2015.
  25. 25. Nanba H, Okumura M. Towards multi-paper summarization reference information. In: Proceedings of the 16th International Joint Conference on Artificial Intelligence, 1999. 926–31.
  26. 26. Teufel S, Carletta J, Moens M. An annotation scheme for discourse-level argumentation in research articles. In: Proceedings of the ninth conference on European chapter of the Association for Computational Linguistics -, 1999. 110. https://doi.org/10.3115/977035.977051
  27. 27. Zhang C, Liu L, Wang Y. Characterizing references from different disciplines: A perspective of citation content analysis. Journal of Informetrics. 2021;15(2):101134.
  28. 28. Wahle JP, Lima Ruas T, Abdalla M, Gipp B, Mohammad SM. Citation amnesia: On the recency bias of NLP and other academic fields. In: Proceedings of the 31st International Conference on Computational Linguistics, 2025. 1027–44. https://aclanthology.org/2025.coling-main.69/
  29. 29. Gustafson DP, Kuehl CR. Citation Age Distributions for Three Areas of Business. J BUS. 1974;47(3):440.
  30. 30. Menéndez ML, Pardo JA, Pardo L, Pardo MC. The Jensen-Shannon divergence. Journal of the Franklin Institute. 1997;334(2):307–18.
  31. 31. Adams J. Early citation counts correlate with accumulated impact. Scientometrics. 2005;63(3):567–81.
  32. 32. Hirsch JE. Does the H index have predictive power?. Proc Natl Acad Sci U S A. 2007;104(49):19193–8. pmid:18040045
  33. 33. Levitt JM, Thelwall M. A combined bibliometric indicator to predict article impact. Information Processing & Management. 2011;47(2):300–8.
  34. 34. Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–74. pmid:843571
  35. 35. Wang D, Song C, Barabási A-L. Quantifying long-term scientific impact. Science. 2013;342(6154):127–32. pmid:24092745
  36. 36. Lee DK. Alternatives to P value: confidence interval and effect size. Korean J Anesthesiol. 2016;69(6):555–62. pmid:27924194
  37. 37. Fanelli D, Glänzel W, Archambault E, Gingras Y, Lariviere V. A bibliometric test of the hierarchy of the sciences: Preliminary results. In: Proceedings of STI 2012 Montreal, 2012. 452–3.
  38. 38. Fanelli D, Glänzel W. Bibliometric evidence for a hierarchy of the sciences. PLoS ONE. 2013;8(6):e66938.
  39. 39. Iqbal S, Hassan S-U, Aljohani NR, Alelyani S, Nawaz R, Bornmann L. A decade of in-text citation analysis based on natural language processing and machine learning techniques: an overview of empirical studies. Scientometrics. 2021;126(8):6551–99.
  40. 40. Zhang Y, Wang Y, Sheng QZ, Yao L, Chen H, Wang K, et al. Deep learning meets bibliometrics: A survey of citation function classification. Journal of Informetrics. 2025;19(1):101608.