Peer Review History
| Original SubmissionOctober 20, 2025 |
|---|
|
-->PONE-D-25-55666-->-->Contrastive Learning with Mutual Information Enhancement and Negative Sample Augmentation Combined with KAN for Text Clustering-->-->PLOS One Dear Dr. Xie, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Apr 08 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript:-->
If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols. We look forward to receiving your revised manuscript. Kind regards, Ning Cai, Ph.D. Section Editor PLOS One Journal Requirements: When submitting your revision, we need you to address these additional requirements. 1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 2. In your Methods section, please include additional information about your dataset and ensure that you have included a statement specifying whether the collection and analysis method complied with the terms and conditions for the source of the data. 3. Thank you for stating the following in the Acknowledgments Section of your manuscript: [This work has been partially supported by Sichuan Science and Technology Program (Grant No: 2023YFQ0044)] We note that you have provided funding information that is not currently declared in your Funding Statement. However, funding information should not appear in the Acknowledgments section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form. Please remove any funding-related text from the manuscript and let us know how you would like to update your Funding Statement. Currently, your Funding Statement reads as follows: "The authors received no specific funding for this work.” Please include your amended statements within your cover letter; we will change the online submission form on your behalf. 4. We note that there is identifying data in the Supporting Information file <GoogleNews-s, GoogleNews- T and GoogleNews-Ts>. Due to the inclusion of these potentially identifying data, we have removed this file from your file inventory. Prior to sharing human research participant data, authors should consult with an ethics committee to ensure data are shared in accordance with participant consent and all applicable local laws. Data sharing should never compromise participant privacy. It is therefore not appropriate to publicly share personally identifiable data on human research participants. The following are examples of data that should not be shared: -Name, initials, physical address -Ages more specific than whole numbers -Internet protocol (IP) address -Specific dates (birth dates, death dates, examination dates, etc.) -Contact information such as phone number or email address -Location data -ID numbers that seem specific (long numbers, include initials, titled “Hospital ID”) rather than random (small numbers in numerical order) Data that are not directly identifying may also be inappropriate to share, as in combination they can become identifying. For example, data collected from a small group of participants, vulnerable populations, or private groups should not be shared if they involve indirect identifiers (such as sex, ethnicity, location, etc.) that may risk the identification of study participants. Additional guidance on preparing raw data for publication can be found in our Data Policy (https://journals.plos.org/plosone/s/data-availability#loc-human-research-participant-data-and-other-sensitive-data) and in the following article: http://www.bmj.com/content/340/bmj.c181.long. Please remove or anonymize all personal information (Name), ensure that the data shared are in accordance with participant consent, and re-upload a fully anonymized data set. Please note that spreadsheet columns with personal information must be removed and not hidden as all hidden columns will appear in the published file. 5. Thank you for stating the following in the Financial Disclosure section: [“The authors received no specific funding for this work.”]. We note that one or more of the authors are employed by a commercial company: name of commercial company. 1. Please provide an amended Funding Statement declaring this commercial affiliation, as well as a statement regarding the Role of Funders in your study. If the funding organization did not play a role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript and only provided financial support in the form of authors' salaries and/or research materials, please review your statements relating to the author contributions, and ensure you have specifically and accurately indicated the role(s) that these authors had in your study. You can update author roles in the Author Contributions section of the online submission form. Please also include the following statement within your amended Funding Statement. “The funder provided support in the form of salaries for authors [insert relevant initials], but did not have any additional role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript. The specific roles of these authors are articulated in the ‘author contributions’ section.” If your commercial affiliation did play a role in your study, please state and explain this role within your updated Funding Statement. 2. Please also provide an updated Competing Interests Statement declaring this commercial affiliation along with any other relevant declarations relating to employment, consultancy, patents, products in development, or marketed products, etc. Within your Competing Interests Statement, please confirm that this commercial affiliation does not alter your adherence to all PLOS ONE policies on sharing data and materials by including the following statement: ""This does not alter our adherence to PLOS ONE policies on sharing data and materials.” (as detailed online in our guide for authors http://journals.plos.org/plosone/s/competing-interests) . If this adherence statement is not accurate and there are restrictions on sharing of data and/or materials, please state these. Please note that we cannot proceed with consideration of your article until this information has been declared. Please include both an updated Funding Statement and Competing Interests Statement in your cover letter. We will change the online submission form on your behalf. 6. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions -->Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. --> Reviewer #1: Yes Reviewer #2: Partly Reviewer #3: Partly Reviewer #4: Yes ********** -->2. Has the statistical analysis been performed appropriately and rigorously? --> Reviewer #1: N/A Reviewer #2: No Reviewer #3: Yes Reviewer #4: Yes ********** -->3. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: No Reviewer #2: Yes Reviewer #3: Yes Reviewer #4: Yes ********** -->4. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.--> Reviewer #1: No Reviewer #2: Yes Reviewer #3: Yes Reviewer #4: Yes ********** -->5. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)--> Reviewer #1: The manuscript presents an experimental evaluation and a technically sophisticated framework. The topic is relevant, and the proposed approach shows potential. However, several issues limit the scientific rigor of the contribution in its current form. Addressing the concerns outlined below is necessary to strengthen the manuscript and improve its credibility. 1. Insufficient Bibliographic Support for Conceptual and Methodological Claims A major concern throughout the manuscript is the lack of adequate bibliographic support for several conceptual and methodological assertions. This issue is particularly evident in: • Statements regarding the limitations of contrastive learning in linear embedding spaces; • General discussions on feature redundancy, anisotropy in embedding spaces, and the curse of dimensionality, which are treated as well-known phenomena but are often not supported by foundational or recent studies; • Assertions that specific mechanisms, such as mutual information maximization, feature-level redundancy reduction, or KAN-based transformations, effectively or significantly mitigate issues, without references to prior theoretical or empirical work motivating these expectations. 2. Overly Assertive Language and Strength of Claims The manuscript frequently employs strong and absolute phrasing (e.g., “effectively solves”, “unequivocally outperforms”, “state-of-the-art performance”). In several instances, these claims are not proportionally supported by statistical significance analysis, controlled comparisons, or evaluations across a sufficiently broad range of experimental conditions. Stronger claims should be reinforced with additional quantitative evidence or references to prior studies reporting similar outcomes. 3. Implicit Technical Assumptions and Lack of Definitions Several technical assumptions are implicitly treated as common knowledge, despite not being universally established. Examples include: • The relationship between mutual information magnitude and semantic dependency at the feature level; • The assumed superiority of learnable spline-based activations over linear transformations in high-dimensional clustering contexts; • The generalization capability of KAN architectures in mitigating the curse of dimensionality. While these assumptions may be plausible, they should be either supported by references or clearly framed as hypotheses validated empirically in this study. In addition, several technical terms (e.g., training steps, topic popularity, false negative samples) are introduced without definition or are defined only after being used. Definitions should be provided at first occurrence. 4. Limited Engagement with Alternative or Conflicting Viewpoints Although the Related Work section is extensive, it primarily emphasizes literature that supports the proposed approach. The manuscript would benefit from a more balanced discussion that also considers: • Alternative paradigms addressing redundancy or anisotropy in embedding spaces; • Potential drawbacks or trade-offs associated with mutual information–based regularization; • Known limitations of contrastive learning frameworks that may also apply to the proposed method. 5. Absence of a Dedicated Discussion Section While the Results section includes localized interpretations of experimental outcomes, the manuscript lacks a dedicated Discussion section. This limits higher-level integration of findings and weakens the connection between empirical results and the broader literature. Although not strictly mandatory, the inclusion of a Discussion section is strongly recommended given the methodological complexity of the proposed framework. 6. Dependence on Hyperparameters and Design Choices The performance of the proposed method appears sensitive to hyperparameter settings (e.g., λ values, batch size, number of clusters). While hyperparameter tuning is mentioned, several architectural and experimental choices, such as encoder selection, KAN configuration, and loss weighting, are not sufficiently justified through references or systematic empirical analysis. Where possible, ablation studies or references to prior work should be used to justify key design decisions. 7. Interpretation of Visual Analyses Visualizations based on techniques such as t-SNE and temporal topic popularity trends are informative, but their interpretations are occasionally too strong. Given the qualitative and parameter-sensitive nature of these methods, their illustrative role should be explicitly acknowledged, and they should not be treated as direct quantitative evidence of superiority. In particular, the concept of “topic popularity” requires a clear quantitative definition. The scale used in the corresponding figures (e.g., 0–1000) should be explicitly explained in the text. Specific Comments • Consider adding a paragraph at the end of the Introduction outlining the structure of the manuscript. • Several paragraphs are overly long. Shorter paragraphs (ideally no more than 10–12 lines) would improve readability. • Figures and tables should be placed closer to their first citation in the text and should be explicitly referenced before appearing. • Some acronyms and abbreviations are not defined at first occurrence. Please ensure all abbreviations are spelled out in full upon first use. Abstract • Lines 13–14: Consider removing the specific days and stating the time frame more generally, for example as “during January and February 2025”. Introduction • Lines 18–20: Additional references are required to support the claims made in this passage. • Lines 25–32: Several key concepts (e.g., contrastive learning, anisotropy, and linear alignment strategies) are introduced without sufficient explanation. Brief contextualization would improve clarity and accessibility for readers less familiar with representation learning. • Line 33: Please spell out the acronym “SCCL” at its first occurrence in the text. This should be consistently verified for all acronyms throughout the manuscript. • Line 76: Consider replacing “pitting” with “evaluating”. • Certain technical terms and claims (e.g., “robust text representations”) could be phrased more precisely or briefly. • Simplifying long sentences, particularly in the paragraph introducing the proposed method, would improve accessibility. 2. Related Work • Line 79: Consider replacing “Related work” with “Related Work” for consistency. • Line 88: Consider using “CNNs”. • Line 158: Consider using “NLP”. • Line 161: Please correct the missing space in “generation.For”. • Line 163: The citation “Pairsupcon” is not numbered and does not appear in the References list. • Splitting long paragraphs and consolidating overlapping descriptions would improve readability. 3. Proposed Model • Line 196: Consider replacing “dubbed” with “termed”. • Line 196: The acronym “MIKAN” is introduced without definition. Please spell out its full meaning at first mention. • Line 201: Consider replacing “ameliorate” with “mitigate”. • Lines 207–208: Consider using only “contrastive learning” to avoid redundancy. • Line 208: Consider replacing “Evidently” with “Consequently” to soften the assertiveness of the statement. • Line 210: Consider rephrasing as “achieving more robust …” for improved clarity. • Line 221: Consider replacing “ameliorate” with “mitigate” for consistency. • The three challenges are currently formulated using “How to …” constructions, which read as implicit questions. Rephrasing these as declarative statements would better suit a methodological section. • Some statements appear overly assertive (e.g., claims that certain mechanisms “fully capture” specific properties). Softening these formulations would help maintain a cautious scientific tone. • Certain theoretical statements, such as those regarding the flexibility of KANs due to learnable activation functions, would benefit from supporting references. • While the overall workflow is clearly presented, briefly clarifying where and how mutual information optimization and feature-level redundancy reduction are implemented within the architecture would further strengthen clarity and reproducibility. 3.1 MIKAN • Lines 224–226: This passage appears redundant and may be removed, as the same definition is already provided in Lines 194–196. • Figure 1: Consider using the label “Pre-trained Encoder” for consistency. In addition, the same shade of green is used at two different stages, which may unintentionally suggest that these elements correspond to the same dataset or representation. 3.2 Mutual Information … • The definition of entropy may need revision. Entropy is typically defined for a random variable rather than a specific realization. Accordingly, the notation should need to be changed from H(x) to H(X) (Formula 1), with the summation taken over the support of the random variable. 3.3 Feature-level Redundancy … • Line 342 and 380: Consider removing the word “Evidently”. • Line 369: Please ensure consistent use of quotation marks for ‘easy’. • The notation appears to be inconsistent throughout the section: batch size is alternately denoted as “b” and “B”, and feature dimensionality as “d” and “D”. Adopting a single notation would improve clarity. • The normalization of the feature matrix may require clarification. Dividing by the sum of squared feature values does not correspond to standard L2 normalization. If the intention is to project features onto the unit hypersphere, normalization would typically involve the square root of the sum of squared elements, which directly affects feature geometry and similarity computation. • The probabilistic interpretation of the self-correlation matrix Mself appears to be unclear. While treating a normalized affinity matrix as an approximation of a joint probability distribution can be acceptable in practice, this assumption should be explicitly stated. As currently written, P(Ei, Ej) seems to be treated as a true probability distribution without explicitly satisfying the axioms of probability. • The definitions of the marginal probability distributions P(Ei) and P(Ej) appear to be ambiguously formatted and do not clearly specify the summation indices. If this interpretation is correct, providing explicit formulations would improve clarity and reproducibility. • The role of the balancing parameter α is only briefly described. A clearer theoretical interpretation of α (e.g., as smoothing or scaling parameter) would strengthen the proposed loss function. 3.4 Negative Sample ... • Equations (13–16): Please correct the typo “Achor” to “Anchor” throughout the equations. • The notation used throughout the section appears to be inconsistent. The total number of samples is denoted interchangeably as “B” and “N”, while additional symbols such as “M”, “L”, and “T” are introduced without clear and unified definitions. • The loss function associated with hard negative samples (Eq. 14) appears difficult to interpret, as it combines regular negatives, hard negatives, and weighting terms. The interaction between the amplified repulsion term and the standard contrastive denominator may require further clarification. • While the attraction mechanism proposed for false negative samples is conceptually reasonable, the formulation of the attraction weight and the corresponding loss functions (Eq. 15 and 16) may benefit from clearer theoretical justification. Although an analogy to attention mechanisms is mentioned, this connection does not appear to be formally established. • The role of the proposed loss LIHSF within the overall optimization framework is not explicitly stated. The authors should clarify whether this loss functions as a primary objective or as an auxiliary regularization term, and how it interacts with the other loss components of the model. 3.5 Kolmogorov–Arnold Networks … • Lines 420–422: The authors are encouraged to add an appropriate reference. • Lines 453–454: Please consider adding a reference. • Line 461: Please ensure consistent use of quotation marks for “curse of dimensionality”. • While the mathematical formulation of the KAN forward propagation and activation functions appears generally correct, claims regarding model efficiency may require clarification. The assertion that KANs substantially reduce the number of parameters compared to fully connected layers is not sufficiently justified. • The statement that the proposed approach alleviates the “curse of dimensionality” appears somewhat strong. While KANs can improve function approximation and representation expressiveness in high-dimensional spaces, a more cautious formulation indicating partial mitigation of high-dimensional effects may be more appropriate. • The computational complexity and scalability of the KAN clustering module are not discussed. Given the use of spline functions and the configuration S = Din, a brief discussion of computational and memory implications would strengthen the section. • While the clustering loss based on pseudo-labels and cross-view consistency is conceptually appropriate, key operational details remain unclear, such as whether pseudo-labels are updated iteratively, whether gradient flow is stopped during pseudo-label generation, and how training stability is ensured. • The rationale for producing three separate score matrices (Y’, Y’1, and Y’2) and their precise role within the overall optimization process should be described more explicitly. 4.2 Experimental Dataset • While the definition of the dataset statistics (N, C, A, and R) is clear, the relevance of the imbalance ratio “R” could be more explicitly discussed. 4.3 Experimental Environment and Parameters • Line 544: Please ensure consistent quotation formatting for the model name. • The reported average runtime per dataset (18 h) should be clarified, for example by indicating whether it includes data preprocessing steps. • When referring to the use of t-SNE for visualization, please clarify whether default parameters were used or specify the key parameter settings adopted. 4.4.1 Comparative Results and Analysis • Line 569: Consider removing “exemplified by BOW and TF-IDF” due to redundancy. • Line 570: Consider removing the word “modern”. • Line 596: Please verify whether the reported value should be “72.69%” instead of “72%”. • Lines 597–598: The source of the reported values is unclear. • Line 611: Consider adding “are unavailable in this study” for clarity. • Line 635: Replace “its” with “MIST” to avoid ambiguity. • Line 635: The origin of the reported value “9.16%” is unclear. • Lines 645–646: Please add an appropriate reference. • Lines 646–648: Please explicitly clarify the stated exception (GoogleNews-TS, SCCL, and NMI). • Line 672: Replace “chapter” with “section”. • Lines 675–678: Please add a supporting reference. • The discussion predominantly emphasizes ACC. A more balanced discussion explicitly referencing NMI values would strengthen the evaluation. • Reported percentage improvements should clearly specify that they represent absolute percentage-point gains to avoid ambiguity. • While the critique of traditional methods (e.g., BOW, TF-IDF, GSDPMM) is technically accurate, the tone is occasionally overly strong. • The discussion would benefit from clearer transitions linking the limitations of existing methods to the motivation and advantages of the proposed approach. • Terminology referring to the BERT baseline should be standardized (e.g., “BERT-based baseline” or “BERT encoder baseline”). • Several sentences—particularly in the discussion of MIST—are long and dense. Splitting these sentences would improve readability. • Explanations for dataset-specific performance degradation (e.g., Biomedical, StackOverflow, Tweet) should be framed more cautiously. • Claims that KAN “effectively mitigates the curse of dimensionality” should be softened. • The statement that the results “demonstrate optimal performance across eight datasets” may be interpreted as overly broad, particularly given that the text itself acknowledges cases where the proposed model does not achieve the best results across all metrics. 4.4.2 Ablation Study ... • Table 5: Please verify the table title. It appears that should be “AgNews, …”. • Tables 5 and 6: The information presented in these tables is redundant with Figures 3 and 4. The authors are encouraged to consider retaining only one type of visualization (tables or figures). • The text refers to “changes in clustering metrics across different datasets”, whereas the ablation study is conducted exclusively on the GoogleNews-TS subset. • The conclusion that the results “conclusively demonstrate the robustness …” is strongly phrased. 4.4.3 Hyperparameter Tuning ... • Line 712: Consider removing the phrase “which significantly … of 1”, as it is redundant. • Line 714: Consider removing the phrase “substantially larger … two datasets”. • The justification for the selected search ranges of λ₁, λ₂, and λ₃ is not explicitly discussed. • The hyperparameter tuning protocol requires clarification. Why were these values selected? • The interpretation of the results is largely descriptive and partially repetitive of Figures 5–7. Condensing the narrative and emphasizing insights would improve conciseness. • The conclusion that λ₁ = 10 and λ₂ = λ₃ = 1×10⁻⁴ are “optimal” should be softened. 4.4.4 Results and Analysis ... • Figure 8: Please consider increasing the font size of numerical values and textual elements. • Statements regarding “O(n)” time complexity and the positive correlation between training time and data scale are conceptually correct but not empirically supported. • The conclusion that the proposed model “necessitates a substantial data scale” to achieve optimal performance should be softened to reflect the experimental scope. 4.4.5 A Case Study ... • Line 759: Please verify the quotation formatting for “DeepSeek”. • Lines 763–764: The phrase “a prominent … platform” appears redundant, as the same information is already stated in Line 759. • Line 764: Replace “chapter” with “subsection”. • Line 769: Please verify the quotation formatting for “elbow point”. • Figure 9: Consider translate the terms into English in Figure 9(a). • Line 779: Replace “The” with “This”. • Lines 780–782: Please verify the quotation formatting for “DeepSeek”, “Innovation and Development”, and “Robotics”. • Lines 782–784: The wording is overly strong. Please moderate terms such as “robust”, “superior”, and “strong”. • Line 796: Consider replacing “curtail” with “reduce”. • Given the qualitative and parameter-sensitive nature of t-SNE, claims regarding model efficacy or stability should emphasize their illustrative role. • The criteria used to assign the chosen semantic labels to clusters are not explicitly stated. • The concept of “topic popularity” is not formally defined. The y-axis scale in Figure 10 (0–1000) should be explained. • The concluding discussion linking topic popularity trends to broader notions such as “new quality productive forces” is somewhat speculative. 5. Conclusion • The conclusion does not explicitly state whether the main research objective of the study has been achieved. • Several performance-related claims are expressed in strong and unqualified terms. While the reported results appear competitive, such statements should be more softened. • The section does not acknowledge any limitations of the proposed approach. • The conclusion should include a discussion of future work. Supplementary Information / Data Availability • The manuscript lists multiple “Supplementary Files” (Files 1–8). However, the provided links appear to redirect only to publicly available / unvailable dataset repositories rather than to supplementary files prepared by the authors. • The manuscript lists eight Supplementary Files, including SearchSnippets and Tweet, but these two datasets are not included in the material provided. Reviewer #2: The manuscript exhibits mixed technical soundness with critical gaps that undermine confidence in the methodology. The most significant technical flaw is the inadequate explanation of the clustering mechanism itself. The paper does not clearly describe how the KAN network transforms embeddings into cluster assignments. Additionally, there is a fundamental conceptual confusion about whether this is truly a clustering method. The approach requires labeled benchmark data for training, yet clustering's primary purpose is to discover structure in unlabeled data. Training on gold-standard labels to learn how to cluster appears contradictory, if labels are available for training, one is performing classification, not unsupervised clustering. The authors do not adequately explain how a method trained on labeled data can generalise to truly unsupervised scenarios like the Weibo dataset, or what advantage this offers over standard supervised classification approaches. I think if this was made clearer then this method has a lot of potential. On the positive side, the individual components appear theoretically grounded. The mutual information enhancement and contrastive learning components are built on established principles from information theory and self-supervised learning. The ablation studies are properly conducted and demonstrate that each proposed component contributes meaningfully to overall performance, which strengthens confidence in the modular design. Regarding whether the data support the conclusions, the evidence is mixed. For the benchmark dataset experiments, there is strong quantitative evidence across eight datasets using standard metrics (ACC and NMI) that supports the claims of superior performance compared to existing baselines. However, the Weibo case study presents significant evidential problems. The conclusions about identifying "DeepSeek," "robotics," and "innovation" topic clusters rely solely on word cloud visualisations and manual interpretation, which constitutes insufficient evidence for these claims. To improve upon this, I would like to see a more rigerous approach to verifying the clusters created. Furthermore, the paper lacks statistical significance testing. Results are presented as point estimates without error bars, confidence intervals, or hypothesis tests. We cannot determine whether the observed improvements over baseline methods represent genuine advances or are within the range of random variance. There is also no mention of multiple experimental runs with different random seeds, which is standard practice for neural network research to account for initialisation sensitivity. In summary, while the benchmark results provide convincing quantitative support for performance claims, the real-world application conclusions are inadequately supported, and the core methodological description requires substantial clarification before the work can be considered technically sound. Reviewer #3: PONE-D-25-55666 Comments Thank you for the opportunity to review the manuscript proposing a text clustering model (MIKAN) that combines mutual information to enhance contrastive learning, negative sample expansion, and KAN network. The research direction has theoretical value and practical application prospects. However, there are many significant issues in the article, including structural integrity, experimental design, literature review, and result analysis. Notably, the expression and support of core novelty, the reproducibility and comprehensiveness of experiments, and the depth of result analysis all need major revision. ABSTRACT 1.The expression of the core novelty point is not precise enough. It merely mentions “combining mutual information enhanced contrastive learning with KAN network” without explicitly explaining the collaborative mechanism among mutual information enhancement, negative sample expansion, and KAN. Meanwhile, it is unclear how this collaborative mechanism specifically addresses the three core issues of text clustering (insufficient representation robustness, redundancy, and dimensionality disaster). You need to supplement the expression with “maximizing mutual information to explore the nonlinear relationship between positive samples, expanding negative samples to optimize sample discrimination, and adapting the KAN network to the complex structure of high-dimensional data, which work together to solve the three core problems of text clustering”, highlighting the innovative logic. 2.The description of the experimental part is overly concise, merely stating that “it performs better on eight datasets” without mentioning the representative characteristics of the datasets (such as data size, domain diversity, category distribution differences, etc.), nor providing specific improvement margins for key performance indicators, thus lacking persuasiveness. INTRODUCTION 1.The research background is not sufficiently systematic, and the development trajectory of text clustering is not clearly outlined. It merely lists the limitations of existing methods without explaining the technological evolution logic behind the mainstream adoption of contrastive learning + deep learning. Furthermore, it fails to identify the common root causes of current research gaps, such as neglecting feature-level redundancy and inadequate modeling of nonlinear relationships. You need to refine the research problem. 2.The discussion on related work such as SCCL, RSTC, and MIST merely states “proposed XXX method” without delving into a comparison of the core flaws and potential improvements of each method, which results in the necessity and entry point of this study not being sufficiently prominent. I suggest that you should strengthen literature comparative discussion and commentary to highlight the necessity of the research. 3.The correspondence between research objectives and novelty is not clear. You merely vaguely mentions “solving three major problems” without explicitly explaining how each novelty point precisely corresponds to a research problem, resulting in a broken logical chain. RELATED WORK 1.The classification system is unreasonable, as it confuses traditional methods, deep learning methods, and contrastive learning methods without establishing a clear classification framework based on technological evolution or method types, leading to a chaotic logic in the literature review. Therefore, please reconstruct the classification framework. Provide a comprehensive review categorized as traditional text clustering methods → deep learning text clustering methods → contrastive learning text clustering methods → high-dimensional data processing methods. For each category, elaborate on the core ideas, representative methods, and limitations, thereby enhancing logical clarity. 2.The overview of positive and negative sample augmentation methods in contrastive learning is not comprehensive enough, omitting important recent research such as the mixed negative sample strategy of MixCSE and the soft negative sample method of SNCSE. Simultaneously, you fail to discuss the applicable scenarios and limitations of different augmentation strategies. Thus, you can add comments on the latest research such as MixCSE constructs difficult negative samples by mixing positive and negative sample features, balancing discrimination and computational efficiency; SNCSE treats the negation form of anchor samples as soft negative samples to alleviate the problem of feature suppression. In addition, you can analyze limitations such as token-level expansion easily destroys semantics, and supervised negative samples rely on labeled data. 3.The review of related research on KAN network is absent, with only a brief mention of KAN in the introduction, and no comprehensive analysis of KAN’s application status in high-dimensional data processing and clustering tasks in the related work section. This results in a lack of literature support for the rationality of introducing this technology. 4.Now, it is 2026 and your references is outdated, as the literature from the past three years are completely insufficient. PROPOSED MODEL 1.The description of the model architecture diagram (Figure 1) is not detailed enough, as it fails to clearly indicate the input and output of each module, data flow, and core parameters. It prevents readers from intuitively understand the model’s operational process. I suggest that refine the description of the model architecture diagram: the original text can be divided into two branches after preprocessing. One branch is subjected to data augmentation and encoder generation to produce M₁⁺ and M₂⁺, which are then output as M₁⁺' and M₂⁺' after passing through the mutual information enhancement module, and subsequently optimized by the IHSF module. The other branch is encoded by the encoder to generate M, which is output as M' after feature redundancy reduction. Finally, the three inputs are fed into the KAN network to obtain clustering results." In addition, label key parameters such as the encoder output dimension D=768 should be highlighted. 2.The derivation and explanation of key formulas are insufficient, such as the absence of the derivation process for mutual information loss functions (Loss_MII, Loss_MIF), the lack of clarity in the physical significance of hyperparameters like α and β, and inconsistent symbol definitions in the formulas (e.g., some symbols are not explained in the main text). 3.Please clarify the module collaboration mechanism. Supplementing the statement like: The outputs M₁⁺' and M₂⁺' of the mutual information enhancement module are used as inputs to the IHSF module, and the optimized features are fused with M' and input into the KAN network; the total loss function balances clustering loss, contrastive loss, and mutual information loss through λ₁, λ₂, and λ₃, achieving end-to-end collaborative optimization. 4.Please supplement KAN configuration details. You need to clarify parameters such as KAN network hidden layer dimension S=768 (consistent with input dimension), using 3rd order B-spline function, learning rate set to 1e-5, adopting L2 regularization (weight decay coefficient 1e-4), and activation function is x·sigmoid(x) combined with spline function, to ensure reproducibility. 5.The Explanation of the basis for weight setting is required. Thus, please supplementing the statement as follows: The optimal values of λ₁ ∈ {0, 1, ..., 100}, λ₂ ∈ {1e-5, ..., 0.1}, and λ₃ ∈ {1e-5, ..., 0.1} were determined through grid search, and the final settings were λ₁=10, λ₂=1e-4, and λ₃=1e-4. This setting achieved the optimal balance between ACC and NMI on the validation set. EXPERIMENTS 1.The comprehensiveness of the experimental design is insufficient, as no control variable experiments such as comparisons of different encoders and different data augmentation strategies beyond ablation experiments were set up. This makes it impossible to verify the independent contributions of each component and the robustness of the model. 2.The description of the dataset is not detailed enough. You only provides basic information such as the number of samples and categories, without specifying the data preprocessing steps (such as text cleaning, word segmentation, stop word removal, etc.). Furthermore, it does not provide key features such as the degree of imbalance in category distribution and text length distribution of the dataset, which limits the reproducibility of the experiment. 3.The selection and setting of baseline models are not reasonable enough. Some baseline models (such as STC²-LPI, Self-Train) have not been tested on all datasets, and the parameter configurations of the baseline models (such as encoder type, training steps, etc.) have not been specified, leading to unfair comparison. 4.The analysis of experimental results is superficial, merely comparing ACC and NMI values without analyzing the performance differences and reasons of the model on different types of datasets (such as balanced/unbalanced, short/long texts), nor does it analyze the rationality of clustering results through visualization (such as intra-cluster compactness, inter-cluster separability). 5.Supplementary efficiency analysis should be added. Add a model efficiency comparison table to compare the parameter quantity (such as MIKAN parameter quantity of 12M and SCCL parameter quantity of 10M), training time (such as MIKAN training for 18 hours and RSTC training for 16 hours on the AgNews dataset), and inference speed (such as the number of samples processed per second) of the model and the baseline model, and analyze the advantages and disadvantages of the model’s efficiency. 6.The description of the collection and preprocessing methods for the case data is not detailed enough, and the crawling strategy for Weibo data (such as the specific scope, frequency, and deduplication method of crawling keywords) is not explained. The parameter configuration and embedding dimensions of the Chinese text embedding model (bge base zh-v1.5) are also not explained, which affects the reproducibility of the case. 7.The basis for determining the number of clusters is insufficient. It merely relies on the elbow method and silhouette coefficient to set k=3, without verifying the clustering results for k=2 and k=4, nor does it incorporate domain knowledge to justify the reasonableness of k=3 (such as whether it aligns with the actual classification logic of scientific topics). 8.There is a problem of insufficient analysis and interpretation of the case results. Your results only displays clustering results through word cloud and T-SNE visualization, without in-depth analysis of the temporal evolution rules of each topic (such as the peak popularity of DeepSeek topics and event correlation), user participation characteristics, etc., nor demonstrating the actual application value of the case results. CONCLUSION 1.The connection between research conclusions and novelty is not close enough, only summarizing the advantages of the model in a general way, without corresponding to the research questions and novelty points raised in the introduction one by one, resulting in an incomplete research logic loop. 2.The explanation of research contributions is not clear enough, without distinguishing between theoretical contributions and practical application contributions, and without highlighting the breakthrough of the research (such as whether KAN has been combined with mutual information enhanced contrastive learning for text clustering for the first time). 3.Your future research directions are too generalized, only mentioning “optimizing model performance” without proposing specific and feasible improvement directions. Some potential directions for your reference: optimizing the clustering performance of imbalanced data through weighted mutual information; extending the model to multi language text clustering, combined with cross language embedding techniques; reducing the number of model parameters and improve real-time inference efficiency. Good luck! Reviewer #4: This manuscript proposes a text clustering approach that integrates contrastive learning, information theory, and a Kolmogorov–Arnold Network (KAN) architecture to address major challenges in text clustering, including the curse of dimensionality and data redundancy. The authors enhance the contrastive learning framework through mutual information maximization strategies, thereby producing more robust text representations while minimizing feature-level redundancy. By increasing the information shared within positive pairs and reducing inter-feature redundancy, the model yields more discriminative representations. Furthermore, by replacing traditional multilayer perceptrons with a KAN architecture, the proposed approach improves its capacity to model complex, high-dimensional data structures. Experimental results on eight datasets demonstrate that the method achieves state-of-the-art performance in terms of accuracy and related evaluation metrics. The experimental design and results section are presented in a detailed and informative manner. The findings are comprehensively compared with existing studies in the literature, and the results are compelling. In Figure 2, the iterative performance improvements of the MIKAN model are visually illustrated. Additionally, a thorough ablation study is conducted to demonstrate the contribution of each component of the model. The methodology for determining the λ hyperparameters is also clearly explained. Points for Improvement: 1.It is recommended that the authors clarify how eliminating correlation redundancy at the feature level differs from traditional dimensionality reduction techniques, and explicitly discuss the conceptual distinctions. 2.The manuscript could be further strengthened by including references to transformer-based pre-trained models applied to text clustering, such as: i.Yasin Ortakci,Revolutionary text clustering: Investigating transfer learning capacity of SBERT models through pooling techniques,Engineering Science and Technology, an International Journal,Volume 55,2024,101730,ISSN 2215-0986,https://doi.org/10.1016/j.jestch.2024.101730. ii. Ortakci, Y., Borhan, B. Optimizing SBERT for long text clustering: two novel approaches with empirical insights. J Supercomput 81, 950 (2025). https://doi.org/10.1007/s11227-025-07414-4 3.On page 4, line 163, a reference citation appears to be missing and should be added. 4.The expression “integrates contrastive learning enhanced by mutual information and negative samples with KAN (Kolmogorov–Arnold Network)” and its variants are repeated multiple times throughout the manuscript. This recurring formulation, used to describe the proposed model, is somewhat redundant and could be streamlined to improve readability. 5.In Section 3.1, the explanation of the proposed MIKAN model based on Figure 1 includes numbered steps (e.g., (1), (2)); however, the chronological ordering and clarity of these steps are insufficient. Rather than enhancing understanding, the current structure tends to create confusion regarding the processing flow. This section would benefit from being reorganized with clearer sequencing and more explicit explanations of each stage. 6.Although the experimental results are detailed, the ACC and NMI scores for the Biomedical dataset are noticeably lower than those for the other datasets. It is recommended that the authors further investigate and discuss the compatibility between the proposed method and the characteristics of the datasets, particularly addressing the reasons behind the relatively lower performance on the Biomedical dataset. Such an analysis would further strengthen the manuscript. ********** -->6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.--> Reviewer #1: Yes: Ismael Weber Reviewer #2: No Reviewer #3: Yes: Zeyu Li Reviewer #4: No ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation. NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications. |
| Revision 1 |
|
Contrastive Learning with Mutual Information Enhancement and Negative Sample Augmentation Combined with KAN for Text Clustering PONE-D-25-55666R1 Dear Dr. Xie, We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. Kind regards, Ning Cai, Ph.D. Section Editor PLOS One Additional Editor Comments (optional): The sole reviewer’s minor revision requests are purely editorial in nature. I see no scientific or technical barriers, and I am confident the authors can resolve them during final production. Accordingly, I approve the manuscript for publication as is. Reviewers' comments: Reviewer's Responses to Questions -->Comments to the Author 1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.--> Reviewer #1: All comments have been addressed Reviewer #3: All comments have been addressed Reviewer #4: All comments have been addressed ********** -->2. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. --> Reviewer #1: Yes Reviewer #3: Yes Reviewer #4: Yes ********** -->3. Has the statistical analysis been performed appropriately and rigorously? --> Reviewer #1: Yes Reviewer #3: Yes Reviewer #4: Yes ********** -->4. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: Yes Reviewer #3: Yes Reviewer #4: Yes ********** -->5. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.--> Reviewer #1: Yes Reviewer #3: Yes Reviewer #4: Yes ********** -->6. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)--> Reviewer #1: I appreciate the considerable effort made by the authors in revising the manuscript. After examining the revised version and the detailed response letter, I am satisfied that the authors have adequately addressed the concerns raised in my previous review. The manuscript has improved substantially in terms of scientific rigor, methodological clarity, bibliographic support, discussion of results, and overall presentation. The authors have strengthened the theoretical foundation of the study, improved the justification of methodological choices, incorporated additional analyses, clarified several technical aspects, and added a dedicated Discussion section. Overall, the revised manuscript presents a stronger and more convincing contribution than the original submission. I have only a few minor comments before publication: 1. Although readability has improved, some sections still contain relatively long and dense paragraphs. A final editorial revision aimed at improving readability would be beneficial. 2. The Abstract could be further strengthened by: * adding a brief introductory sentence to better contextualize the research problem; * highlighting one or two key quantitative results obtained by the proposed method; * explicitly stating the main outcome of the Weibo case study validation rather than only mentioning that the validation was performed. 3. There appears to be an incomplete placeholder in line 983 (“peak at k = 3 (0.XX), indicating…”), which should be corrected before publication. These issues are primarily editorial in nature and do not affect my overall positive assessment of the manuscript. Therefore, I recommend Minor Revision. Reviewer #3: Thank you for your detailed response. I think that you have addressed all of the concerns. Good luck. Reviewer #4: I believe that the newly added sections in the discussion have made the article clearer, more coherent, and more compelling. Additionally, the revision of Section 3.1 has successfully resolved the previously existing inconsistencies. ********** -->7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.--> Reviewer #1: Yes: Ismael Weber Reviewer #3: Yes: Zeyu Li Reviewer #4: No ********** |
| Formally Accepted |
|
PONE-D-25-55666R1 PLOS One Dear Dr. Xie, I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team. At this stage, our production department will prepare your paper for publication. This includes ensuring the following: * All references, tables, and figures are properly cited * All relevant supporting information is included in the manuscript submission, * There are no issues that prevent the paper from being properly typeset You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps. Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing. If we can help with anything else, please email us at customercare@plos.org. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Ning Cai Section Editor PLOS One |
Open letter on the publication of peer review reports
PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.
We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.
Learn more at ASAPbio .