Peer Review History

Original SubmissionMay 18, 2026
Decision Letter - Bilal Alatas, Editor

Dear Dr. Mhagama,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Sep 14 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

  • A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.
  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.
  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

As the corresponding author, your ORCID iD is verified in the submission system and will appear in the published article. PLOS supports the use of ORCID, and we encourage all coauthors to register for an ORCID iD and use it as well. Please encourage your coauthors to verify their ORCID iD within the submission system before final acceptance, as unverified ORCID iDs will not appear in the published article. Only the individual author can complete the verification step; PLOS staff cannot verify ORCID iDs on behalf of authors.

We look forward to receiving your revised manuscript.

Kind regards,

Bilal Alatas, Ph.D.

Academic Editor

PLOS One

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please note that PLOS One has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

3. We note you have included a table to which you do not refer in the text of your manuscript. Please ensure that you refer to Table 1 in your text; if accepted, production will need this reference to link the reader to the Table

4. Please ensure that you refer to Figures 4, 6 and 11 in your text as, if accepted, production will need this reference to link the reader to the figure.

5. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

Additional Editor Comments:

Dear Authors,

Feedback from the reviewers is now available. It is not recommended that your article be published in its current format. However, we strongly recommend that you address the issues raised by the reviewers, especially those related to readability, methodology, experimental design and validity, and resubmit your paper after making the necessary changes.

Best wishes,

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

Reviewer #1: Partly

Reviewer #2: Partly

**********

2. Has the statistical analysis been performed appropriately and rigorously? -->?>

Reviewer #1: N/A

Reviewer #2: No

**********

3. Have the authors made all data underlying the findings in their manuscript fully available??>

The PLOS Data policy

Reviewer #1: No

Reviewer #2: No

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: No

Reviewer #2: Yes

**********

Reviewer #1: This paper considers precision agriculture, combining computer vision (leaf image analysis) with environmental context (weather data) to improve crop disease diagnosis and generate farmer-facing recommendations via a LLM, using a multimodal fusion approach.

The experimental dataset is small and imbalanced, the multimodal fusion architecture lacks technical implementation specifics, the LLM evaluation is absent, and several textual typos and layout inconsistencies must be resolved.

1. Some of the references in the reference list, are with missing data (eg. Ref [3])

2. Introduction – Refer more latest related work from 2026, 2025 and clearly state how your contribution covers the latest gaps.

3. In the introduction, state the research questions addressed in this study. In the discussion section, describe how you fulfilled the said research questions, using the proposed methodology and the obtained results.

4. Related studies – it would be nice to discuss the techniques used in latest studies such as https://doi.org/10.1109/SCSE70081.2026.11499862

5. The image dataset contains 661 healthy samples but only 135 for Maize Streak Virus and 77 for Aphids. Because the data is highly imbalanced, relying primarily on global Accuracy (~94%) can be highly misleading. Therefore, provide class-specific metrics (Precision, Recall, and F1-Score) for all three models in Table 3.

6. Why didn’t you use any data augmentation or class-weighted loss functions during training to protect the minority disease classes from being ignored by the optimizer. Is it possible to run the model again and get the results by addressing this.

7. The environmental data spans only 61 days (August to September 2022). Can you state, exactly how a single daily weather observation vector was paired with individual images. If multiple images were taken on the same day, did they share identical weather vectors? If yes, discuss the risk of data leakage or artificial correlation during the random 80:20 train/test split.

8. Multimodal Fusion Specifics: In Section 3.6, you state that image embeddings from MobileNetV2 and weather features are concatenated. Generally, MobileNetV2 outputs high-dimensional embeddings , and the weather vector consists of just 4 variables (temperature, humidity, rainfall, solar radiation). Simple concatenation would allow the image features to completely overwhelm the weather features numerically. Specify if any feature scaling, dimensionality reduction, or projection layers were applied to the text/image vectors before concatenation.

9. In this study, a Random Forest was used for the weather-only model, while deep neural layers were used for the fusion model. Can you justify this selection in technical terms. Why a simple MLP was not used for the weather-only model. Explain how the Random Forest outputs were seamlessly integrated if late fusion was considered instead of early concatenation.

10. Please provide evaluation metrics for LLM-generated recommendataions in Section 4.4, showing numerical validation of the text output quality.

11. Is it possible to include human expert evaluation (e.g., scoring by Agronomists on a 1–5 scale for Safety, Technical Accuracy, and Actionability) or automated language metrics (like BERTScore or G-Eval) to guarantee the system does not produce agricultural hallucinations.

12. Include a comparison table to compare the proposed results with the existing latest studies.

13. What is the possibility of deploying this model in real-world as in https://doi.org/10.1109/SCSE70081.2026.11499862

14. Clearly specify the exact GPT model version utilized and state the core hyperparameter settings used during inference.

15. Proofread the paper for typos.

16. It would be better to provide the data availability link for the 61-day environmental dataset CSV file to an open repository (such as Zenodo alongside the YEESI dataset link) .

Reviewer #2: The manuscript addresses a timely and important topic by proposing a multimodal agricultural decision support system that integrates image-based disease detection, meteorological data, and recommendation generation. The study has practical application potential, and the overall organization of the article (introduction, methods, results, and discussion) is clear and well-structured. However, significant shortcomings in methodology, experimental validation, and reporting must be addressed before the article can be published.

The most significant issue is that the experimental validation supporting the superiority of the proposed multimodal approach is not sufficiently robust. An accuracy of approximately 91% is reported for the image-based model and approximately 94% for the multimodal model. However, it has not been demonstrated whether this performance difference is statistically significant. Confidence intervals, p-values, or appropriate statistical comparisons have not been provided. Furthermore, the evaluation was conducted using only a single 80%/20% training-test split, and neither k-fold cross-validation nor an independent validation dataset was used.

The dataset used in the study is relatively small and unbalanced across classes. In particular, the number of samples in the aphid class is quite low. In addition, the meteorological data covers only 61 days of observations and a single geographic region. This limits the model’s generalizability to different regions and different growing conditions. It is recommended that these limitations be discussed in greater detail and, if possible, supported by additional validation experiments.

The technical details of the multimodal fusion model should be explained more comprehensively. In particular, the feature dimensions used, the fusion method, hyperparameters, learning rate, batch size, number of epochs, data augmentation strategies, and the training process should be provided in detail. This information is necessary for the reproducibility of the study.

Although the recommendation engine was presented as one of the study’s significant contributions, it has not been sufficiently evaluated. The version of the GPT model used, the prompt design, and how the accuracy of the recommendations was verified were not explained. An evaluation by agricultural experts or a user-centered performance analysis would significantly strengthen the study’s scientific contribution.

It is also recommended that the study include comparisons with stronger baseline methods and present an ablation study demonstrating the contribution of each component of the proposed architecture. These analyses would more convincingly demonstrate the extent to which the multimodal approach actually contributes.

There are also shortcomings regarding data accessibility. Although access to the image dataset was provided, the meteorological dataset, the final data structure used for multimodal fusion, training codes, model weights, and the prompts for the recommendation system were not shared. Therefore, it does not appear possible to fully reproduce the study.

Finally, the English text of the article should undergo careful editing. There are numerous grammatical and terminological errors throughout the text (e.g., “Fusion model” instead of "Fusion modal" "Weather-Based Model" instead of "Weather-Base Modal," and "plant pasts" instead of "plant pests"). These corrections will improve the article’s readability and academic quality.

In conclusion, the study presents an interesting idea with high potential for application. However, it requires strengthening of methodological details, more comprehensive experimental validation, the inclusion of statistical analyses, improvements in data and code sharing, and a thorough linguistic revision.

My overall recommendation: Major Revision.

**********

what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

Revision 1

REVIEWER I

1 Some of the references in the reference list are missing bibliographic information (e.g., Ref. [3]).

We thank the reviewer for identifying this issue. The reference list has been carefully reviewed and corrected. Missing bibliographic information, including complete titles, journal names, publication years, volume/issue numbers, page or article numbers, and DOIs (where available), has been added to comply with the journal's reference formatting requirements.

2 Introduction – Refer more latest related work from 2026, 2025 and clearly state how your contribution covers the latest gaps.

We sincerely thank the reviewer for this valuable suggestion. The Introduction has been substantially revised to incorporate recent studies published in 2025 and 2026 covering deep learning, multimodal learning, explainable AI, and large language models for plant disease diagnosis. The revised literature review also identifies a clear gap in the latest research: although recent approaches increasingly integrate multiple data modalities, most continue to focus primarily on disease detection and classification, with limited integration of context-aware recommendation generation and quantitative evaluation of AI-generated agricultural advice. The revised Introduction explicitly positions our contribution within this gap by proposing a multimodal context-aware AI recommender framework that integrates image-based disease diagnosis, environmental information, multimodal feature fusion, rule-based agronomic knowledge, and LLM-assisted recommendation generation, followed by blind LLM-based evaluation of recommendation quality. These revisions clarify the novelty and relevance of the proposed framework in relation to the latest research.

3 In the introduction, state the research questions addressed in this study. In the discussion section, describe how you fulfilled the said research questions, using the proposed methodology and the obtained results.

We appreciate the reviewer's valuable suggestion. The Introduction has been revised to explicitly state the research questions that guided this study. In addition, the Discussion section has been expanded to explain how each research question was addressed through the proposed multimodal methodology and validated using the experimental results. These revisions improve the clarity of the study objectives and strengthen the connection between the methodology, results, and conclusions.

4 Related studies – it would be nice to discuss the techniques used in latest studies such as https://doi.org/10.1109/SCSE70081.2026.11499862

We thank the reviewer for this valuable suggestion. The Related Work section has been revised to include recent studies published in 2025–2026, including the suggested work. We now discuss the underlying techniques employed in these studies, such as CNN-based disease classification with Grad-CAM explainability, multimodal learning, and large language model integration. Furthermore, we contrast these approaches with our proposed framework by highlighting that, unlike existing studies, our work combines multimodal disease classification with a hybrid recommendation engine that generates context-aware, actionable recommendations for farmers.

5 The image dataset contains 661 healthy samples but only 135 for Maize Streak Virus and 77 for Aphids. Because the data is highly imbalanced, relying primarily on global Accuracy (~94%) can be highly misleading. Therefore, provide class-specific metrics (Precision, Recall, and F1-Score) for all three models in Table 3.

We thank the reviewer for this valuable observation. The manuscript has been revised by expanding Table 3 to include class-specific Precision, Recall, and F1-score for the weather-based, image-based, and multimodal fusion models. These additional metrics provide a more informative assessment of model performance, particularly for the minority disease classes.

It is important to note that the dataset was preserved in its original class distribution because it represents the natural occurrence of healthy and diseased maize plants under real agricultural field conditions. Rather than artificially balancing the dataset, which may overestimate real-world performance, our objective was to evaluate the proposed framework under realistic deployment scenarios.

The manuscript has been revised accordingly.

6 Why didn’t you use any data augmentation or class-weighted loss functions during training to protect the minority disease classes from being ignored by the optimizer. Is it possible to run the model again and get the results by addressing this.

We sincerely thank the reviewer for this valuable suggestion. Following the recommendation, we conducted additional experiments to investigate the effects of data augmentation, class-weighted loss, and their combination on the naturally imbalanced maize disease dataset. These experiments were performed using the same MobileNetV2 architecture, training configuration, and evaluation protocol to ensure a fair comparison.

The results indicated that while these imbalance mitigation techniques influenced class-specific performance, they did not improve the overall balance between classification accuracy and minority-class recognition for our real-world dataset. Specifically, data augmentation reduced the overall classification accuracy from 90% to 85%, while class-weighted training further reduced the accuracy to approximately 70% by introducing excessive false-positive predictions for the minority classes. The combination of augmentation and class-weighted loss exhibited similar behavior.

These findings suggest that the naturally imbalanced dataset reflects the actual distribution of maize diseases observed under field conditions in Morogoro, Tanzania. Consequently, aggressive imbalance mitigation strategies caused the model to overcompensate for minority classes, reducing its ability to correctly recognize the dominant healthy class and lowering the overall generalization performance.

Based on these comparative experiments, we retained the original image model because it provided the best trade-off between overall accuracy and class-specific performance. The manuscript has been revised to include class-specific Precision, Recall, and F1-score metrics together with an additional discussion of these experiments.

7 The environmental data spans only 61 days (August to September 2022). Can you state exactly how a single daily weather observation vector was paired with individual images? If multiple images were taken on the same day, did they share identical weather vectors? If yes, discuss the risk of data leakage or artificial correlation during the random 80:20 train/test split.

We thank the reviewer for this insightful observation. The environmental dataset consisted of daily weather observations obtained from the NASA POWER database for the Morogoro study area between August and September 2022. Each maize leaf image was associated with the weather measurements corresponding to its date of acquisition. Consequently, multiple images collected on the same day shared the same environmental feature vector while retaining their own image data and disease labels.

We acknowledge that this alignment may introduce correlation among samples collected on the same day. Because the image and environmental datasets were randomly divided into training and testing subsets using an 80:20 split, it is possible that images acquired on the same date appeared in both subsets and therefore shared identical environmental attributes. This could potentially introduce a degree of information leakage through the environmental modality.

However, the environmental variables constitute only one component of the proposed multimodal framework. The primary discriminatory information is derived from the maize leaf images, while the environmental variables provide complementary contextual information rather than unique identifiers of disease classes. This is supported by the independent evaluation of the weather-only model, which achieved substantially lower performance (74% accuracy; macro F1-score = 0.34) than both the image-based model and the multimodal fusion model. These results indicate that the environmental features alone have limited predictive capability and primarily serve to enhance, rather than dominate, the classification process.

To clarify this aspect, the manuscript has been revised by explicitly describing the image-weather alignment procedure and by acknowledging this potential limitation in the Discussion section.

8 Multimodal Fusion Specifics: In Section 3.6, you state that image embeddings from MobileNetV2 and weather features are concatenated. Generally, MobileNetV2 outputs high-dimensional embeddings, and the weather vector consists of just 4 variables (temperature, humidity, rainfall, solar radiation). Simple concatenation would allow the image features to completely overwhelm the weather features numerically. Specify if any feature scaling, dimensionality reduction, or projection layers were applied to the text/image vectors before concatenation.

We thank the reviewer for this important observation. In the proposed framework, the weather variables (temperature, relative humidity, rainfall, and solar radiation) were standardized using the StandardScaler, where the scaler was fitted on the training set and subsequently applied to the testing set to prevent information leakage.

Rather than directly concatenating the original high-dimensional MobileNetV2 features, the proposed framework extracts a compact 128-dimensional image embedding from the trained image model. This embedding is then concatenated with the four standardized weather variables, producing a 132-dimensional multimodal feature vector.

The fused representation is subsequently processed through fully connected dense layers (128 and 64 neurons), which serve as learnable projection layers that automatically optimize the relative contribution of both modalities during end-to-end training. Consequently, feature scaling together with the compact embedding representation prevents numerical dominance of the image modality and enables effective integration of visual and environmental information.

The manuscript has been revised to clarify these implementation details.

9 In this study, a Random Forest was used for the weather-only model, while deep neural layers were used for the fusion model. Can you justify this selection in technical terms. Why a simple MLP was not used for the weather-only model. Explain how the Random Forest outputs were seamlessly integrated if late fusion was considered instead of early concatenation.

We sincerely thank the reviewer for this important observation. We agree that the methodology description could have been clearer.

The Random Forest model was not integrated into the multimodal fusion network. Instead, it was developed as an independent baseline model to evaluate the predictive capability of environmental variables alone. The purpose of this model was to quantify the discriminative power of weather information without visual features and to provide a baseline for comparison with the image-based and multimodal approaches.

The proposed multimodal framework employs early feature-level fusion, rather than late decision-level fusion. Specifically, deep image embeddings extracted from the trained MobileNetV2 model were concatenated directly with the standardized environmental variables (temperature, relative humidity, rainfall, and solar radiation). The resulting multimodal feature vector was subsequently processed by fully connected neural network layers to perform disease classification. Consequently, the outputs of the Random Forest classifier were not used during multimodal fusion.

Regarding the selection of Random Forest for the environmental baseline, the environmental dataset consisted of only four numerical variables and a relatively limited number of observations. Random Forest was selected because of its robustness for structured tabular data, ability to capture nonlinear relationships, resistance to overfitting on small datasets, and interpretability through feature importance analysis. In contrast, multilayer perceptrons generally require substantially larger datasets to achieve stable generalization and would have introduced unnecessary model complexity for the weather-only baseline.

To avoid confusion, Section 3.6 has been revised to explicitly distinguish the independent weather baseline from the proposed multimodal fusion framework and to clarify that the fusion strategy is based on early feature concatenation rather than late fusion.

10 Please provide evaluation metrics for LLM-generated recommendations in Section 4.4, showing numerical validation of the text output quality.

Thank you for this valuable suggestion. We have extended Section 4.5 by introducing a quantitative evaluation of the generated recommendations. Thirty representative recommendation cases covering all disease classes were evaluated using an independent Large Language Model (Llama 3.1) under a blind assessment protocol.

The evaluator received only the predicted disease, environmental conditions, and generated recommendation without access to the underlying agronomic rules. Recommendation quality was assessed using five criteria: Safety, Technical Accuracy, Relevance, Actionability, and Clarity on a five-point Likert scale.

The resulting quantitative evaluation has been incorporated into the revised manuscript as Table 4 and discussed in Section 4.5.

11 Is it possible to include human expert evaluation (e.g., scoring by Agronomists on a 1–5 scale for Safety, Technical Accuracy, and Actionability) or automated language metrics (like BERTScore or G-Eval) to guarantee the system does not produce agricultural hallucinations?

Thank you for the insightful recommendation. Human expert evaluation by agronomists would indeed provide the strongest external validation. However, access to multiple certified agronomists was beyond the scope of the present study. Instead, following recent literature on LLM evaluation, we introduced an automated blind evaluation protocol using Llama 3.1 as an independent evaluator.

The evaluator scored recommendation quality across Safety, Technical Accuracy, Relevance, Actionability, and Clarity without access to the original agronomic rules. This evaluation provides quantitative evidence that the generated recommendations preserve agronomic correctness while maintaining high readability. We acknowledge that future work will incorporate human expert evaluation by agricultural specialists to further validate recommendation quality under real farming conditions.

12 Include a comparison table to compare the proposed results with the existing latest studies.

Thank you for this valuable suggestion. We have addressed this comment by adding a new comparison table (Table 6) in Section 4.6, comparing the proposed framework with four recent state-of-the-art studies published in 2026. The comparison includes the methodology, dataset, data modality, recommendation capability, classification performance, and key contribution of each study. Unlike previous approaches that primarily focus on disease classification or diagnostic reporting, the proposed framework integrates multimodal disease classification with context-aware recommendation generation and quantitative evaluation of AI-generated agricultural recommendations.

A more comprehensive experimental comparison of the proposed approach against existing methods using standardized performance metrics is part of the next objective of this research, which focuses on evaluating the proposed model against existing approaches. The present study therefore provides a state-of-the-art comparison based on the performance metrics reported in the selected recent studies, while the broader comparative evaluation will be undertaken as part of the subsequent research objective.

13 What is the possibility of deploying this model in real-world as in https://doi.org/10.1109/SCSE70081.2026.11499862?

Thank you for the suggestion. We have clarified the real-world deployment potential of the proposed framework. The system is intended to be deployed as a mobile application where farmers can capture maize leaf images, retrieve weather information automatically or enter it manually, and receive real-time disease diagnosis together with context-aware management recommendations. This deployment perspective has been added as future work under discussion section in the revised manuscript

14 Clearly specify the exact GPT model version utilized and state the core hyperparameter settings used during inference.

Thank you for highlighting the importance of reproducibility. We have revised the manuscript to explicitly specify the large language model used for recommendation generation and evaluation. The revised manuscript states that Llama 3.1 (8B), deployed locally through Ollama, was used. This information has been added to the implementation details in the methodology.

15 Proofread the paper for typos.

Thank you for the suggestion. The manuscript has been carefully proofread and revised to correct typographical, grammatical, punctuation, and formatting errors. We also improved sentence clarity, standardized terminology, and ensured consistency in figures, tables, and references throughout the manuscript.

16 It would be better to provide the data availability link for the 61-day environmental dataset CSV file to an open repository (such as Zenodo alongside the YEESI dataset link).

Thank you for the valuable suggestion. The dataset used in this study are publicly available at https://doi.org/10.5281/zenodo.7729284 and the weather dataset and codes are available at https://github.com/mhagama/research

REVEWER II

SN Reviewer Comment Response

1 An accuracy of approximately 91% is reported for the image-based model and approximately 94% for the multimodal model. However, it has not been demonstrated whether this performance difference is statistically significant. Confidence intervals, p-values, or appropriate statistical comparisons have not been provided.

We appreciate the reviewer's valuable observation. The reported improvement from the image-based model (91%) to the proposed multimodal model (94%) demonstrates a consistent performance gain across all evaluation metrics, including precision, recall, and F1-score. We acknowledge that the original manuscript did not include statistical significance testing or confidence interval analysis.

This has now been explicitly recognized as a limitation of the current study in the revised manuscript. Future work will include formal statistical validation, such as confidence interval estimation and hypothesis testing, using larger multi-site datasets to further quantify the statistical significance of the observed performance improvements.

2 The evaluation was conducted using only a single 80%/20% training-test split, and neither k-fold cross-validation nor an independent validation dataset was used.

We thank the reviewer for this valuable comment. A stratified 80%/20% train-test split was deliberately adopted to ensure a fair comparison by evaluating the image-based, weather-based, and multimodal models on the same data partition while preserving class distribution. Since the primary objective was to assess the contribution of multimodal fusion under identical experimental conditions, this evaluation protocol was considered appropriate. We acknowledge that k-fold cross-validation and independent external validation would further strengthen the generalizability analysis and have identified these as directions for future work.

3 The dataset used in the study is relatively small and unbalanced across classes. In particular, the number of samples in the aphid class is quite low. In addition, the meteorological data covers only 61 days of observations and a single geographic region. This limits the model’s generalizability to different regions and different growing conditions. It is recommended that these limitations be discussed in greater detail and, if possible, supported by additional validation experiments.

We thank the reviewer for this important observation. We acknowledge that the dataset is relatively small, exhibits class imbalance (particularly for the aphid class), and that the environmental observations cover only a 61-day period from a single geographical region.

These characteristics reflect the natural field conditions under which the data were collected and were intentionally preserved to evaluate the proposed framework under realistic deployment scenarios rather than artificially balanced conditions. To address this concern, the Discussion section has been expanded to explicitly describe these limitations, their potential impact on model generalizability, and future plans to validate the framework using larger, balanced, multi-season, and multi-location datasets.

4 The technical details of the multimodal fusion model should be explained more comprehensively. In particular, the feature dimensions used, the fusion method, hyperparameters, learning rate, batch size, number of epochs, data augmentation strategies, and the training process should be provided in detail. This information is necessary for the reproducibility of the study.

We thank the reviewer for this valuable suggestion. The methodology has been substantially revised to improve the reproducibility of the proposed framework. The revised manuscript now provides detailed information on the feature dimensions, multimodal fusion strategy, image embedding extraction, weather feature preprocessing, network architecture, hyperparameter settings (learning rate, batch size, number of epochs, optimizer), training procedure, and data augmentation techniques. These additions provide a comprehensive description of the multimodal fusion model and facilitate reproducibility of the proposed approach.

5 Although the recommendation engine was presented as one of the study’s significant contributions, it has not been sufficiently evaluated. The version of the GPT model used, the prompt design, and how the accuracy of the recommendations was verified were not explained. An evaluation by agricultural experts or a user-centered performance analysis would significantly strengthen the study’s scientific contribution.

We thank the reviewer for this valuable comment. The recommendation engine has been substantially expanded in the revised manuscript. A dedicated evaluation section has been added describing the recommendation generation process, the LLM version and inference settings, prompt design, and a quantitative blind evaluation framework based on Safety, Technical Accuracy, Relevance, Actionability, and Clarity. The evaluation demonstrates that the generated recommendations are context-aware, agronomically consistent, and grounded in the rule-based recommendation engine. We agree that evaluation by professional agronomists and user-centered field studies would further strengthen the practical validation of the recommendation system and have identified these as important directions for future work.

6 It is also recommended that the study include comparisons with stronger baseline methods and present an ablation study demonstrating the contribution of each component of the proposed architecture. These analyses would more convincingly demonstrate the extent to which the multimodal approach actually contributes.

We thank the reviewer for this valuable suggestion. The revised manuscript has been strengthened by including additional comparisons with recent state-of-the-art studies published in 2026. Furthermore, the experimental evaluation includes separate image-based, weather-based, and multimodal models, all trained and evaluated under identical experimental conditions. These comparative experiments demonstrate the contribution of each modality and the performance improvement achieved through multimodal feature fusion. We acknowledge that a more comprehensive component-wise ablation study, including systematic evaluation of individual environmental variables and alternative fusion strategies, would provide further insight into the contribution of each component.

A more comprehensive experimental comparison of the proposed approach against existing methods using standardized performance metrics is part of the next objective of this research, which focuses on evaluating the proposed model against existing approaches. The present study therefore provides a state-of-the-art comparison based on the performance metrics reported in selected recent studies, while the broader comparative and component-wise evaluation will be undertaken as part of the subsequent research objective.

7 There are also shortcomings regarding data accessibility. Although access to the image dataset was provided, the meteorological dataset, the final data structure used for multimodal fusion, training codes, model weights, and the prompts for the recommendation system were not shared. Therefore, it does not appear possible to fully reproduce the study.

Thank you for the valuable suggestion. The dataset used in this study are publicly available at https://doi.org/10.5281/zenodo.7729284 and the weather dataset and codes are available at https://github.com/mhagama/research

8 Finally, the English text of the article should undergo careful editing. There are numerous grammatical and terminological errors throughout the text (e.g., “Fusion model” instead of "Fusion modal" "Weather-Based Model" instead of "Weather-Base Modal," and "plant pasts" instead of "plant pests"). These corrections will improve the article’s readability and academic quality.

Thank you for the suggestion. The manuscript has been carefully proofread and revised to correct typographical, grammatical, punctuation, and formatting errors. We also improved sentence clarity, standardized terminology, and ensured consistency in figures, tables, and references throughout the manuscript.

Attachments
Attachment
Submitted filename: Reviewer Comment.docx
Decision Letter - Bilal Alatas, Editor

A Multimodal Context-Aware AI Recommender for Smart Farming

PONE-D-26-24624R1

Dear Dr. Mhagama,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Bilal Alatas, Ph.D.

Academic Editor

PLOS One

Additional Editor Comments (optional):

Dear Author,

One of the previous reviewers has not sent the review results. However, after carefully cheking your paper seems sufficiently improved and ready for publication.

Best wishes,

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

Reviewer #1: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions??>

Reviewer #1: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously? -->?>

Reviewer #1: N/A

**********

4. Have the authors made all data underlying the findings in their manuscript fully available??>

The PLOS Data policy

Reviewer #1: No

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English??>

Reviewer #1: Yes

**********

Reviewer #1: Paper is improved.

However, proofread the paper well for further improvement of the clarity of the description.

**********

what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy

Reviewer #1: No

**********

Formally Accepted
Acceptance Letter - Bilal Alatas, Editor

PONE-D-26-24624R1

PLOS One

Dear Dr. Mhagama,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS One and supporting open access.

Kind regards,

PLOS One Editorial Office Staff

on behalf of

Prof. Dr. Bilal Alatas

Academic Editor

PLOS One

Open letter on the publication of peer review reports

PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.

We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.

Learn more at ASAPbio .