Peer Review History

Original SubmissionDecember 17, 2024
Decision Letter - Carlos Carrasco-Farré, Editor

PONE-D-24-58179Evaluating the capacity of large language models to interpret emotions in imagesPLOS ONE

Dear Dr. Alrasheed,

Thank you for submitting your manuscript to <em data-end="199" data-start="189">PLOS ONE</em>. We appreciate the effort and thought that went into this research, and we are pleased to inform you that we would like to invite you to submit a revised version of your paper. Based on the reviewers’ assessments, we are offering a revise and resubmit with minor revisions.

Both reviewers acknowledge the relevance and significance of your study, particularly in evaluating GPT-4’s capabilities in recognizing and rating emotions from visual stimuli. Below is a summary of their key comments that I encourage you to address in your revision:

<h3 data-end="781" data-start="753">Reviewer Comments:</h3>

  1. Importance of Research Scope: Reviewer 1 highlighted that your study aligns with applied research trends in emotion recognition and its potential broad applicability. Reviewer 2 suggested strengthening the introduction by explicitly explaining why evaluating emotional cues from non-facial images is important.
  2. Clarity in Numeric Scale Tables: Reviewer 2 noted that the numeric scale tables could be difficult to interpret because they assume prior knowledge about differences between LLM and human ratings. They suggest including a comparison metric to indicate what constitutes a “close” comparison to human ratings and providing a brief explanation of its significance.
  3. Contextual Emotion Recognition Comparison: The manuscript focuses more on facial emotion recognition, but Reviewer 2 suggests discussing contextual emotion recognition and referencing other studies that have used the same dataset, even if they do not involve LLMs.
  4. Figure Placement & Labeling: The placement of figures at the end of the paper makes it difficult for readers to understand their context. Reviewer 2 recommends integrating figures into the relevant sections where they are mentioned in the text and improving figure labels to clarify their significance.
  5. Demographic Details of Human Raters: Given that perceived emotional effects can vary across demographics, Reviewer 2 recommends providing more details on the demographics of human raters to address potential biases in the study.

<h3 data-end="2350" data-start="2311">Additional Editorial Request:</h3>

We also noticed that the image quality in the manuscript is not optimal, as they appear blurred in the current submission. Please ensure that all figures are of high resolution and check that they remain clear after uploading the PDF version.

Overall, your study makes an important contribution to the field, and we are confident that these refinements will enhance its clarity and impact. We look forward to receiving your revised manuscript and thank you for your efforts in addressing these minor revisions.

Congratulations on your work, and we appreciate your contribution to <em data-end="2955" data-start="2945">PLOS ONE</em>.

Please submit your revised manuscript by Apr 11 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.
  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.
  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Carlos Carrasco-Farré

Academic Editor

PLOS ONE

Journal Requirements:

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2.  Please ensure that you refer to Figure 4 in your text as, if accepted, production will need this reference to link the reader to the figure.

3. We notice that your supplementary figures are uploaded with the file type 'Figure'. Please amend the file type to 'Supporting Information'. Please ensure that each Supporting Information file has a legend listed in the manuscript after the references list.

4. We note you have included a table to which you do not refer in the text of your manuscript. Please ensure that you refer to Table 8 and 9 in your text; if accepted, production will need this reference to link the reader to the Table.

5. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Evaluating GPT-4's capability to recognize and rate emotions from visual stimuli is a critical area of research, as it has the potential to become one of the preferred methods in applied emotional and psychological studies. This evaluation is important for several reasons:

+ GPT-4's ability to interpret visual stimuli, such as facial expressions, can provide deeper insights into emotional dynamics across various contexts, from individual well-being to social interactions.

+ The use of GPT-4 for emotion recognition aligns with trends in applied research methodologies, offering scalable, efficient, and adaptable tools for large-scale studies in fields like marketing, sociology, education, and behavioral science.

+ Rigorous testing and evaluation ensure the reliability of GPT-4 in academic and applied research, encouraging broader adoption in scientific communities.

By focusing on these aspects, the evaluation of GPT-4's emotion recognition capabilities could establish it as a versatile and trusted tool in understanding human emotions through visual stimuli.

Reviewer #2: Summary:

* This research is attempting to evaluate the effectiveness of Chat-GPT4o in predicting the emotional response different images have on the subject through both visual and written depiction of images in order to determine if Chat-GPT is as good or better than previous work on explicitly engineered emotional evaluation of direct subject observance using deep learning and deterministic models.

Strengths:

* The paper did a good job of providing multiple evaluation criteria for their central thesis

* I was happy to see all the visual representations of the data and its comparisons within the paper

* I felt the paper made clear it’s conclusion and did a good job supporting that conclusion by highlighting the relevant data

Weaknesses:

* The paper did not establish why evaluation of emotional queues from non-facial images was important.

* The tables on the numeric scale are difficult to evaluate because it assumes the reader understands what difference between the LLM and the human raters is low versus high. 

* The image based related work seems to focus more on facial emotion recognition and not contextual emotion recognition. A well known and highly available dataset is used in this work for this purpose, however it is not compared with existing work. Are there any relevant work related to the same dataset being used? Even if they are not LLMs?

* Other than the tables, each figure is pasted at the end of the paper and the labels of the figures do not provide appropriate context to show what each represents.

Suggestions for strengthening the paper:

* Include a paragraph of why this research matters and how it can benefit others or be expanded to practical applications.

* Include a comparison metric in the numeric scale tables that indicate what is a “close” comparison to the human rating and then briefly explain how the comparison metric is used in the paper.

* Include at least one mention of a study done with the same dataset and evaluation criteria that was used in your research (i.e. likert scale emotion detection)

* It would help the reader better understand the context of each figure in the paper if the figure was shown in the same area as it was mentioned in the paper. Pasting them all at the end makes it very difficult to determine what each figure is intended to show as the labeling does not provide that information. The reader must go into the paper and find where each figure is mentioned.

* There is no mention of what demographic the human raters are. Since this deals with perceived emotional effects of specific images, having a skewed demographic can influence the ratings. Thus, it would be effective to have a diverse demographic participating.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Revision 1

Dear Dr. Carrasco-Farré,

Thank you for the opportunity to submit a revised version of our manuscript, Evaluating the Capacity of Large Language Models to Interpret Emotions in Images, to PLOS ONE. We sincerely appreciate the time and effort that you and the reviewers have dedicated to providing thoughtful and constructive feedback.

We have carefully addressed the reviewers' major concerns and incorporated their suggestions throughout the manuscript to improve clarity and readability. In addition, we have addressed the journal’s formatting and submission requirements as outlined. Specifically:

• We enhanced the quality of the figures to ensure greater clarity.

• We added in-text references to all figures and tables.

• We changed the file type of the supplementary figures to comply with submission guidelines.

• We reviewed and updated the reference list by removing two duplicate entries and several references that were not cited in the current version of the manuscript. Additional references were included in response to the reviewers’ suggestions. We also revised the formatting of all entries to align with PLOS ONE style.

We have also addressed all reviewer comments and suggestions. A detailed, point-by-point response is provided below. The corresponding revisions are clearly marked in red in the tracked version of the manuscript.

Thank you again for considering our resubmission. We look forward to your response.

Sincerely,

Hend Alrasheed

Media Lab, Massachusetts Institute of Technology

hrasheed@mit.edu

April 4, 2025

Point-by-point response to the reviewers’ comments and concerns.

Response to Reviewer 1 Comments

We thank the reviewer for their reading of the manuscript and for their constructive comments.

Comment 1.1: Evaluating GPT-4's capability to recognize and rate emotions from visual stimuli is a critical area of research, as it has the potential to become one of the preferred methods in applied emotional and psychological studies. This evaluation is important for several reasons:

- GPT-4's ability to interpret visual stimuli, such as facial expressions, can provide deeper insights into emotional dynamics across various contexts, from individual well-being to social interactions.

- The use of GPT-4 for emotion recognition aligns with trends in applied research methodologies, offering scalable, efficient, and adaptable tools for large-scale studies in fields like marketing, sociology, education, and behavioural science.

- Rigorous testing and evaluation ensure the reliability of GPT-4 in academic and applied research, encouraging broader adoption in scientific communities.

By focusing on these aspects, the evaluation of GPT-4's emotion recognition capabilities could establish it as a versatile and trusted tool in understanding human emotions through visual stimuli.

Response 1.1: We appreciate the reviewer’s thoughtful comments highlighting the significance of evaluating GPT-4’s ability to recognize and rate emotions from visual stimuli.

Response to Reviewer 2 Comments

We thank the reviewer for their reading of the manuscript and for their constructive comments. We have taken the comments into consideration to improve the quality of the manuscript. Please find below a point by point response to each comment.

Comment 2.1: The paper did not establish why evaluation of emotional queues from non-facial images was important.

Response 2.1: We have added a paragraph to the Introduction that clarifies the importance of evaluating emotional cues from non-facial images (lines: 149-161).

“While most efforts in the literature focus on evaluating the capabilities of LLMs in extracting emotions from facial images, our work focus to emotion recognition from general, non-facial images, such as objects, environments, animals, and abstract scenes. Despite their rich emotional content and widespread use in psychological, affective computing, and mental health research \cite{Geneva,ortis2020survey}, the interpretation of emotions elicited by non-facial imagery remains relatively underexplored, particularly in the context of Large Language Models. Assessing emotional responses to such stimuli is crucial, as they provide opportunities to study affective processing in broader, more ecologically valid contexts where facial expressions may be absent or irrelevant. Moreover, non-facial images are a foundational component of standardized emotional elicitation datasets such as GAPED and IAPS, highlighting their significance in emotion research.”

Comment 2.2: The tables on the numeric scale are difficult to evaluate because it assumes the reader understands what difference between the LLM and the human raters is low versus high.

Include a comparison metric in the numeric scale tables that indicate what is a “close” comparison to the human rating and then briefly explain how the comparison metric is used in the paper.

Response 2.2: We have added a paragraph to the Results section clarifying how the difference between GPT-4 and human ratings should be interpreted using the Mean Absolute Error (MAE) metric. The new text explains the MAE values relative to the 0–100 rating scale (lines: 356-360).

“Note that the maximum possible deviation between GPT-4 and the human rating for any image is 100, as both valence and arousal were rated on a 0–100 scale. In our results, MAE values ranged from approximately 5 to 15, indicating that the average error across all image categories was relatively small, representing only 5–15\% of the maximum possible error.”

Comment 2.3: The image based related work seems to focus more on facial emotion recognition and not contextual emotion recognition. A well known and highly available dataset is used in this work for this purpose, however it is not compared with existing work. Are there any relevant work related to the same dataset being used? Even if they are not LLMs?

Response 2.3: We have added three studies that utilize the GAPED image dataset in distinct emotion elicitation tasks to the Related Work section (lines: 151-158).

“While most efforts in the literature focus on evaluating the capabilities of LLMs in extracting emotions from facial images, our work shifts the focus to emotion recognition from general, non-facial images. Accordingly, we use the GAPED image dataset, a well-established resource for emotion elicitation. Several studies have employed this dataset for a range of emotional research tasks. For example, Moyal et al. \cite{moyal2018Categorized} used GAPED images to elicit discrete emotions such as fear, disgust, sadness, and happiness. Balsamo et al. \cite{Balsamo2020bottom} used valence and arousal ratings to evaluate the effectiveness of GAPED images in evoking emotional responses. Moreover, the author in \cite{Brainerd2018emotional} used GAPED images to explore how different combinations of valence and arousal influence cognitive processing, particularly in emotionally ambiguous contexts.”

Comment 2.4: Other than the tables, each figure is pasted at the end of the paper and the labels of the figures do not provide appropriate context to show what each represents. It would help the reader better understand the context of each figure in the paper if the figure was shown in the same area as it was mentioned in the paper. Pasting them all at the end makes it very difficult to determine what each figure is intended to show as the labeling does not provide that information. The reader must go into the paper and find where each figure is mentioned.

Response 2.4: We apologize for the inconvenience. In accordance with PLOS submission requirements, all figures have been uploaded as separate files and not embedded within the manuscript. We understand that figure links and placements will be handled during the final production stage.

Comment 2.5: Include a paragraph of why this research matters and how it can benefit others or be expanded to practical applications.

Response 2.5: We have added a paragraph to the end of the Introduction that highlights the importance of our work and its potential applications. (lines: 82-86).

“By demonstrating that GPT-4 can closely approximate human emotional ratings of visual stimuli, this work offers a scalable and efficient alternative to traditional emotion validation methods, which are often labor-intensive and costly. Such automation can streamline experimental design in psychology and facilitate the creation of emotionally intelligent AI agents.”

Comment 2.6: Include at least one mention of a study done with the same dataset and evaluation criteria that was used in your research (i.e. likert scale emotion detection).

Response 2.6: We have added a reference to a study that uses the same dataset to elicit valence and arousal ratings using a Likert scale, and have incorporated it to the Related Work section (lines: 158-161).

“A recent study \cite{Berezina2024} used the GAPED dataset to assess valence and arousal ratings within a Malaysian population using a 9-point Likert scale, with the goal of identifying culturally specific patterns in emotional responses.”

Comment 2.7: There is no mention of what demographic the human raters are. Since this deals with perceived emotional effects of specific images, having a skewed demographic can influence the ratings. Thus, it would be effective to have a diverse demographic participating.

Response 2.7: Thank you for this important observation. The demographic information of the human raters is provided in the manuscript in the Image Dataset section, based on the original GAPED dataset documentation. While we rely on GAPED’s existing ratings, we agree that future work should explore expanding demographic diversity to further enhance generalizability.

Attachments
Attachment
Submitted filename: Response to reviewers.pdf
Decision Letter - Carlos Carrasco-Farré, Editor

Evaluating the capacity of large language models to interpret emotions in images

PONE-D-24-58179R1

Dear Dr. Alrasheed,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. If you have any questions relating to publication charges, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Carlos Carrasco-Farré

Academic Editor

PLOS ONE

Formally Accepted
Acceptance Letter - Carlos Carrasco-Farré, Editor

PONE-D-24-58179R1

PLOS ONE

Dear Dr. Alrasheed,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Carlos Carrasco-Farré

Academic Editor

PLOS ONE

Open letter on the publication of peer review reports

PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.

We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.

Learn more at ASAPbio .