Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Development and evaluation of AI model with deep learning for segmentation of extraocular muscles in thyroid eye disease

  • Yusuke Haruna ,

    Roles Data curation, Investigation, Project administration, Resources, Software, Writing – original draft

    mizuki1979feb@yahoo.co.jp

    Affiliation Department of Ophthalmology and Visual Sciences, Graduate School of Medicine, Osaka Metropolitan University, Abeno-ku, Osaka-shi, Osaka-fu, Japan

  • Mizuki Tagami,

    Roles Conceptualization, Formal analysis, Funding acquisition, Investigation, Writing – review & editing

    Affiliation Department of Ophthalmology and Visual Sciences, Graduate School of Medicine, Osaka Metropolitan University, Abeno-ku, Osaka-shi, Osaka-fu, Japan

  • Ryo Kurosaki,

    Roles Investigation

    Affiliation Department of Medical Device Engineering, Graduate School of Medicine, Kobe University, Chuo-ku, Kobe-shi, Hyogo Prefecture, Japan

  • Mizuho Nishio,

    Roles Conceptualization, Investigation, Software, Supervision, Writing – review & editing

    Affiliations Department of Medical Device Engineering, Graduate School of Medicine, Kobe University, Chuo-ku, Kobe-shi, Hyogo Prefecture, Japan, Departments of Radiology, Graduate School of Medicine, Kobe University, Chuo-ku, Kobe-shi, Hyogo Prefecture, Japan

  • Gen Kinari,

    Roles Investigation

    Affiliation Department of Ophthalmology and Visual Sciences, Graduate School of Medicine, Osaka Metropolitan University, Abeno-ku, Osaka-shi, Osaka-fu, Japan

  • Mami Tomita,

    Roles Investigation

    Affiliation Department of Ophthalmology and Visual Sciences, Graduate School of Medicine, Osaka Metropolitan University, Abeno-ku, Osaka-shi, Osaka-fu, Japan

  • Norihiko Misawa,

    Roles Investigation, Supervision

    Affiliations Department of Ophthalmology and Visual Sciences, Graduate School of Medicine, Osaka Metropolitan University, Abeno-ku, Osaka-shi, Osaka-fu, Japan, Department of Ophthalmology & Visual Sciences, Washington University School of Medicine, St. Louis, Missouri, United States of America

  • Yoshihiro Muragaki,

    Roles Supervision, Writing – review & editing

    Affiliation Department of Medical Device Engineering, Graduate School of Medicine, Kobe University, Chuo-ku, Kobe-shi, Hyogo Prefecture, Japan

  • Shigeru Honda

    Roles Supervision, Writing – review & editing

    Affiliation Department of Ophthalmology and Visual Sciences, Graduate School of Medicine, Osaka Metropolitan University, Abeno-ku, Osaka-shi, Osaka-fu, Japan

Abstract

Purpose

To develop and evaluate an AI model for the segmentation of extraocular muscles (EOMs) using Magnetic Resonance Imaging (MRI).

Design

Single-center study with retrospective study.

Methods

The study included 52 patients with thyroid eye disease (TED) who underwent MRI examination of the orbital region at the Department of Ophthalmology, Osaka Metropolitan University, between October 2020 and June 2023. Manual labelling of all EOMs was performed on all slices. An AI model was created and compared with the manually labelled data using 5-fold cross validation on data sets of 12, 32, and 52 cases. The average signal intensity ratio (SIR) within the EOMs was measured. The primary outcome was the comparison of DICE similarity coefficients (DSC) as the agreement rate between AI’s result and manual labelling amongst the three data sets (52 cases vs 32 cases vs 12 cases). The secondary outcome was the correlation of SIR between the AI and manual labelling.

Results

A significant difference in DSC was observed between the 12-, 32-, and 52-case data sets for the inferior rectus muscle and medial rectus muscle. A significant correlation was observed between the AI and manual labelling for SIR in all EOMs.

Conclusions

An AI model was successfully developed for the automatic segmentation of EOMs. There was little measurement error between AI and manual labelling, and the AI model using 52 cases had improved measurement accuracy compared to the AI models using 32 and 12 cases.

Introduction

Thyroid eye disease (TED) is an autoimmune disorder that primarily affects the orbit and periorbital tissues. It is characterized by inflammation and enlargement of the extraocular muscles (EOMs) and orbital fat, which can lead to various ocular symptoms including proptosis, diplopia, and vision loss [1,2]. Therefore, accurate evaluation of EOMs is crucial for quantifying disease activity, monitoring treatment response, and predicting outcomes in TED patients.

In recent years, Artificial Intelligence (AI) has seen significant advancements and has been increasingly applied to various areas within the medical field, particularly in medical imaging [3,4]. Magnetic Resonance Imaging (MRI) plays a vital role in the assessment of TED, offering detailed visualization of orbital structures and being useful for quantifying disease activity and predicting the response to anti-inflammatory treatment and the outcome in TED [5,6]. MRI is a powerful tool for evaluating EOMs involvement and disease activity in TED, as it can detect inflammatory edema before irreversible fibrosis occurs [7]. Despite the growing interest in AI applications in medicine [4], while there is some AI research using CT to diagnose and evaluate TED [8], AI research using MRI is still limited. CT is a useful test, but it is an invasive test that involves radiation exposure, whereas MRI noninvasively generates cross-sectional images of internal structures and allows for anatomical and inflammation assessment [9].

Currently, there is no clear standard for judging the presence or absence of activity in EOMs inflammation using MRI in the diagnosis and treatment of TED. To address this, we focused on the signal intensity ratio (SIR) in Short Tau/Time Inversion Recovery (STIR) sequence as an objective indicator for evaluating intraorbital inflammation. SIR on STIR or TIRM sequences has been reported as a promising objective biomarker that correlates with the Clinical Activity Score (CAS) and predicts treatment response [10]. Additionally, advanced techniques such as diffusion-weighted imaging have been explored to further objectify the ‘tip of the iceberg’ in orbital inflammation [11]. The primary objective of this study was to investigate the feasibility of developing an AI model capable of automatic segmentation of EOMs, which is a foundational step towards developing an AI model for evaluating the presence or absence of disease activity based on SIR. Contrast-enhanced T1-weighted images (T1WI) are the most accurate for volume segmentation, but assessing inflammation using only T1WI is difficult [5]. While previous studies have highlighted the value of combining multiple sequences, such as T1, T2, and STIR, for a comprehensive assessment of disease activity [7,10], our study focused on a single-sequence (STIR) AI model to simplify the clinical workflow. This development of automated methods for segmenting and analyzing EOMs in TED could significantly improve diagnostic efficiency and precision.

Methods

Prior to this study, we obtained approval from the Ethics Committees of the Osaka Metropolitan University (OMU) (approval no. 2023−072 First Approval: 26/09/2023, Second change approval: 31/03/2025, Third change approval:10/07/2025), Japan. All patients at OMU who participated in the study provided optout informed consent for their patient information to be stored in the hospital database and used in this study. This study was conducted in accordance with the tenets of the Declaration of Helsinki.

The study protocol was a single-center study and the protocol was as follows. This study included 52 patients diagnosed with moderate to severe TED who underwent plain MRI of the orbital region at the Department of Ophthalmology, OMU, between October 2020 and June 2023. The diagnosed severity was based on the European Thyroid Association/European Group on Graves’ Orbitopathy (EUGOGO) management guidelines [12]. All patients underwent MRI before treatment. MRI scans were performed using a coronal STIR sequence. MRI scans were performed using a 3.0T scanner [SIEMENS, MAGNETOM Vida] with a HeadNeck_20_TCS. STIR images were acquired in the coronal plane with the following parameters: repetition time (TR) = 4970ms or 5320ms or 5670ms, echo time (TE) = 59ms, and inversion time (TI) = 230ms. The slice thickness was 3 mm with a 0.3 mm interslice gap. The field of view was 180 × 180 mm, and the matrix size was 512 × 512. We accessed the medical records from 30/08/2024–20/12/2024 and exported them to a CD-R as a DICOM image series. Manual labelling of the EOMs—superior rectus muscle (SRM), inferior rectus muscle (IRM), medial rectus muscle (MRM), and lateral rectus muscle (LRM) —was performed on all identifiable slices of both eyes using the image analysis software ITK-snap. (Fig 1) ITK-snap is a segmentation tool widely used in radiology and oncology that allows for the creation of 3D models thru manual labelling [13]. Three ophthalmologist (YH, MT, and NM) performed the labelling. A board-certified radiologist (MN) visually confirmed the results of the labelling by the ophthalmologists. The average SIR within the EOMs was also measured. Additionally, signal intensity within the eyeball was measured by automatic segmentation of the eyeball using Hough transform. The SIR was defined as follows:

thumbnail
Fig 1. Manual labeling EOMs on MRI (STIR sequence) using ITK-snap.

Labeling was done on the EOMs as shown in a (coronal image). b is a transverse images reconstructed from coronal images using ITK-snap. green: SRM, yellow: IRM, red: MRM, blue: LRM.

https://doi.org/10.1371/journal.pone.0349074.g001

Signal intensity ratio (SIR)= EOM signal intensity /eyeball signal intensity

The segmentation model used was U-Mamba_Enc [14,15]. U-Mamba is a model that incorporates the Mamba architecture, which is an extension of State Space Models (SSMs), into nnU-Nett V2 [16,17]. nnU-Nett V2 is a semantic segmentation model specialized for medical images that is based on U-Nett, a Convolutional Neural Network (CNN)-based segmentation model, and has the ability to automatically adapt to a given dataset [18]. By using Selective SSMs, Mamba can capture long-distance dependencies that are difficult to capture with the local receptive field of CNN. Since U-Mamba is a hybrid structure of CNN and Mamba, it was proposed as a method that can efficiently extract both local features and long-distance dependencies. There are two types of U-Mamba: U-Mamba_Enc, which uses U-Mamba blocks in all encoder blocks, and U-Mamba_Bot, which uses U-Mamba blocks only in the bottleneck. In this study, U-Mamba_Enc was used. The model input data was 3D STIR image data, and the ground truth labels were data that specified the area of the EOMs on STIR images at the pixel level. We used the default image preprocessing and default self-configurations that U-Mamba inherits from nnU-Nett V2. The performance of the final model was determined by the average score calculated by 5-fold cross-validation. We used the implementation of 5-fold cross-validation included in U-Mamba. An overview of 5-fold cross-validation is as follows. The input data for the model was divided equally into five folds. One fold was used for validation, and the remaining four folds were used for model training. Five patterns (5 splits) were prepared, with the test fold covering all folds. Regression model training and testing were performed for each split. The evaluation scores for all folds were averaged to determine the final evaluation score. (S1 Appendix).

The primary outcome was the comparison of the agreement rate between the AI model’s results and manual labelling across the 12, 32, and 52-case datasets. The secondary outcome was to evaluate the correlation of SIR between the AI and manual groups across the 52-case datasets.

Statistical analysis

All statistical analyses were performed with EZR (Version 1.60, Saitama Medical Centre, Jichi Medical University, Saitama, Japan), which is used for the calculation of R. More precisely, it is a modified version of R commander, which is designed to add statistical functions frequently used in biostatistics [19]. The Dice similarity coefficient (DSC) was used to quantify the accuracy of image segmentation as agreement rate between AI’s results and manual labelling. One-way ANOVA and Tukey multiple comparisons test was employed for comparisons involving three or more groups. Spearman’s rank correlation coefficient and Bland-Altman plots were used to assess the correlation between SIR measurements. Values of P < 0.05 were considered statistically significant.

Results

An example of the AI’s result is shown in Fig 1B. The DSC for each EOM is shown in Fig 2. For SRM, the DSC was 0.74 ± 0.07 in the 12 group, 0.75 ± 0.06 in the 32 group, and 0.75 ± 0.06 in the 52 group, with no significant differences observed between the groups. (One-way ANOVA: p = 0.70) For IRM, the DSC was 0.72 ± 0.22 in the 12 group, 0.82 ± 0.07 in the 32 group, and 0.82 ± 0.06 in the 52 group, with significant differences observed between the groups. (One-way ANOVA: p = 0.011) For MRM, the DSC was 0.78 ± 0.08 in the 12 group, 0.83 ± 0.04 in the 32 group, and 0.83 ± 0.05 in the 52 group, with significant differences observed between the groups. (One-way ANOVA: p = 0.0092) For LRM, the DSC was 0.69 ± 0.22 in the 12 group, 0.78 ± 0.10 in the 32 group, and 0.78 ± 0.10 in the 52 group, with no significant differences observed between the groups. (One-way ANOVA: p = 0.087) Turkey’s multiple comparison test was performed on the IRM and MRM. For IRM, it found significant differences between 12 group and 32 group (p = 0.020) and between 12 group and 52 group (p = 0.0094). There was no significant difference between 32 group and 52 group (p = 0.99). For MRM, it found significant differences between 12 group and 32 group (p = 0.0084) and between 12 group and 52 group (p = 0.015). There was no significant difference between 32 group and 52 group (p = 0.85).

thumbnail
Fig 2. Concordance rate (DICE coefficient) between AI and manual EOMs labeling in each group (12 group vs 32 group vs 52 group).

Significant differences between groups were observed in IRM and MRM. a: SRM, b: IRM, c: MRM, d: LRM.

https://doi.org/10.1371/journal.pone.0349074.g002

Fig 3 shows the Bland-Altman plots of SIR for the AI and manual measurements using data from 52 cases. The mean SIR for manual and AI measurements was 0.37 ± 0.14 for SRM, 0.36 ± 0.15 for IRM, 0.38 ± 0.12 for MRM, and 0.33 ± 0.077 for LRM. The difference in SIR between manual and AI measurements (manual-AI) was −0.0080 ± 0.020 for SRM, −0.012 ± 0.017 for IRM, −0.0075 ± 0.016 for MRM, and −0.016 ± 0.013 for LRM.

thumbnail
Fig 3. Bland-Altman plots of SIR for the AI and manual measurements using data from 52 cases.

The difference between the AI and Manual measurements was minor. a: SRM, b: IRM, c: MRM, d: LRM.

https://doi.org/10.1371/journal.pone.0349074.g003

Fig 4 shows the correlation between AI and manually measured SIR using data from 52 cases. Spearman’s correlation coefficients were 0.987 for SRM (p = 0), 0.97 for IRM (p = 0), 0.985 for MRM (p = 0), and 0.988 for LRM (p = 0), and significant strong correlations were observed for all EOMs.

thumbnail
Fig 4. Spearman’s correlation coefficients between AI and manually measured SIR using data from 52 cases.

Significant strong correlations were observed for all EOMs. a: SRM, b: IRM, c: MRM, d: LRM.

https://doi.org/10.1371/journal.pone.0349074.g004

Fig 5 shows the representative case of AI model’s result for segmentation of EOMs. Fig 1 and Fig 5(a,b) show the same slice from the same patient. Comparing Fig 1,5, there is a slight difference with SRM (Fig 5c,d: area surrounded by yellow oval). However, for the other EOMs, the AI’ results were almost identical to the manual labelling. We found that the AI’s result was almost as accurate as the manual labelling.

thumbnail
Fig 5. AI’s result of segmenting EOMs on MRI.

Fig 1 corresponds to Fig 5(a, b). C is an enlarged view of the left orbital area of Fig 5a, and d is an enlarged view of the left orbital area of Fig 1a. green: SRM, yellow: IRM, red: MRM, blue: LRM.

https://doi.org/10.1371/journal.pone.0349074.g005

Discussion

In this study, we successfully developed an AI model for the automatic segmentation of EOMs in patients with TED using MRI images. A significant difference in DSC was observed between the 12, 32, and 52-case groups for the IRM and MRM. This suggests that incorporating a larger sample size improved the performance of the AI model, particularly for these muscles. This finding is consistent with the general principle of deep learning that more data often leads to improved model performance [20]. The SRM showed no significant improvement in DSC with increased sample size. The lack of significant difference in the SRM may be attributed to the anatomical proximity and the shared fascial connections between the SRM and the levator palpebrae superioris muscle. In TED, these muscles are often collectively referred to as the Superior Rectus-Levator Palpebrae Superioris (SRLPS) complex due to their tendency to enlarge together and the frequent disappearance of the intervening fat plane on MRI. This anatomical characteristic likely led to partial volume effects and increased the difficulty of precise discrete segmentation, which may have affected the statistical results for the SRM. As mentioned in the introduction, using T1WI may result in more accurate segmentation. However, because MRI images cannot be captured under different conditions simultaneously, slight misalignments in muscle positions can easily occur between STIR and T1 images due to eye movements or slight head movements during imaging. Therefore, when analyzing T1 and STIR images together with a single AI model, “co-registration” is necessary to precisely overlay the two images at the pixel level. STIR images, while less accurate than T1 images, can be segmented, thus avoiding the risk of registration errors that occur in the process of combining multiple images. Furthermore, by loading only one image series, the clinical workflow can be simplified, making segmentation tools easier to use in clinical settings. Also, STIR images are superior for activity assessment in TED [5,21]. So, we decided to use STIR images for segmentation. In contrast, IRM and MRM are known to be the most commonly swollen EOMs in TED [22,23]. This distinct swelling may make them easier to differentiate from surrounding tissues, potentially leading to higher accuracy in manual labelling and, consequently, better training data for the AI model. DSC was 0.75 for SRM, 0.82 for IRM, 0.83 for MRM, 0.78 for LRM in an AI model using 52 cases. A similar study using TED computed tomography (CT) images reported a DSC of 0.902, which was higher than the present study [8]. This may be due to the large sample size of 281 and the fact that muscle tissue boundaries are less clear in the MRI (STIR sequence) compared to CT. In previous reports on AI-based automatic segmentation of the uterus or prostate using T2-weighted MRI, the DSC was 0.927 [24], 0.793 [24], 0.82 [25], which was close to the results of this study.

Furthermore, a significant correlation was found between the AI group and the manual group in terms of the average SIR for all EOMs. The Bland-Altman analysis indicated minimal measurement errors between the two groups. While there are currently no established reference values for SIR, making it difficult to precisely evaluate the magnitude of our study’s measurement error, previous reports comparing SIR before and after treatment in moderate to severe TED showed a decrease from 2.2 ± 1.4 to 1.8 ± 1.4 [26]. Given this background, it appears that the measurement errors observed in this study were minor and that the AI model is capable of accurately measuring the SIR of EOM. In this study, we evaluated the signal intensity of EOMs using the SIR, providing a quantitative basis for identifying ‘hyperintensity’ associated with active TED. While clinical scores (CAS) or healthy controls were not included in this dataset, the AI-derived SIR successfully represented a continuous spectrum of signal intensities. Higher SIR values extracted by the model correspond to the hyperintense areas typically indicative of active inflammation, as described in previous studies [10]. This demonstrates that our AI model can objectively ‘inform’ clinicians of muscle signal status without the inter-observer variability of visual inspection. Future research incorporating clinical activity data will be necessary to define the optimal SIR (for example, thresholds generally considered to be high signal intensity. normal signal intensity or hyperintensity) threshold for active TED. This study is the first step in objectively evaluating various MRI data in the future.

This study has several limitations. First, it was a single-center study. Therefore, it is necessary to verify whether the AI model developed in this study can be used on MRI at other facilities in the future. Second, we only included patients with active TED. It remains to be verified whether the AI model can segment EOMs with similar performance in MRI without EOM swelling after treatment. Third, this segmentation is performed using only STIR images, so the volume estimate may differ slightly from the actual volume. We are trying to create an AI model that evaluates inflammation and volume using only one continuous image, so we chose STIR images, which allow assessment of both inflammation and volume.

Conclusion

We successfully created an AI model capable of automatically segmenting EOMs from MRI images with the STIR sequence. This AI model has the potential to be a useful tool for real world TED practice in the future.

Supporting information

S1 Appendix. Five-fold cross-validation used in the current study.

The input data for the model was divided equally into five folds.(model 1–5) The evaluation scores for all folds were averaged to determine the final evaluation score.

https://doi.org/10.1371/journal.pone.0349074.s001

(DOCX)

References

  1. 1. Douglas RS, Gupta S. The pathophysiology of thyroid eye disease: implications for immunotherapy. Curr Opin Ophthalmol. 2011;22(5):385–90. pmid:21730841
  2. 2. Stan MN, Garrity JA, Bahn RS. The evaluation and treatment of graves ophthalmopathy. Med Clin North Am. 2012;96(2):311–28. pmid:22443978
  3. 3. Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, et al. A survey on deep learning in medical image analysis. Med Image Anal. 2017;42:60–88. pmid:28778026
  4. 4. Chen X, Wang X, Zhang K, Fung K-M, Thai TC, Moore K, et al. Recent advances and clinical applications of deep learning in medical image analysis. Med Image Anal. 2022;79:102444. pmid:35472844
  5. 5. Mayer EJ, Fox DL, Herdman G, Hsuan J, Kabala J, Goddard P, et al. Signal intensity, clinical activity and cross-sectional areas on MRI scans in thyroid eye disease. Eur J Radiol. 2005;56(1):20–4. pmid:15896938
  6. 6. Vaishnav YJ, Mawn LA. Magnetic Resonance Imaging in the Management of Thyroid Eye Disease: A Systematic Review. Ophthalmic Plast Reconstr Surg. 2023;39(6S):S81–91. pmid:38054988
  7. 7. Luccas R, Riguetto CM, Alves M, Zantut-Wittmann DE, Reis F. Computed tomography and magnetic resonance imaging approaches to Graves’ ophthalmopathy: a narrative review. Front Endocrinol (Lausanne). 2024;14:1277961. pmid:38260158
  8. 8. Alkhadrawi AM, Lin LY, Langarica SA, Kim K, Ha SK, Lee NG, et al. Deep-learning based automated segmentation and quantitative volumetric analysis of orbital muscle and fat for diagnosis of thyroid eye disease. Investigative Ophthalmology & Visual Science. 2024;65(5):6. pmid:38696188
  9. 9. Yousaf T, Dervenoulas G, Politis M. Advances in MRI Methodology. Int Rev Neurobiol. 2018;141:31–76. pmid:30314602
  10. 10. Kirsch EC, Kaim AH, De Oliveira MG, von Arx G. Correlation of signal intensity ratio on orbital MRI-TIRM and clinical activity score as a possible predictor of therapy response in Graves’ orbitopathy--a pilot study at 1.5 T. Neuroradiology. 2010;52(2):91–7. pmid:19756565
  11. 11. Politi LS, Godi C, Cammarata G, Ambrosi A, Iadanza A, Lanzi R, et al. Magnetic resonance imaging with diffusion-weighted imaging in the evaluation of thyroid-associated orbitopathy: getting below the tip of the iceberg. Eur Radiol. 2014;24(5):1118–26. pmid:24519110
  12. 12. Bartalena L, Kahaly GJ, Baldeschi L, Dayan CM, Eckstein A, Marcocci C, et al. The 2021 European Group on Graves’ orbitopathy (EUGOGO) clinical practice guidelines for the medical management of Graves’ orbitopathy. Eur J Endocrinol. 2021;185(4):G43–67. pmid:34297684
  13. 13. Yushkevich PA, Piven J, Hazlett HC, Smith RG, Ho S, Gee JC, et al. User-guided 3D active contour segmentation of anatomical structures: significantly improved efficiency and reliability. Neuroimage. 2006;31(3):1116–28. pmid:16545965
  14. 14. Ma J, Li F, Wang B. U-mamba: Enhancing long-range dependency for biomedical image segmentation. arXiv preprint. 2024.
  15. 15. https://github.com/bowang-lab/U-Mamba
  16. 16. Isensee F, Jaeger PF, Kohl SAA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat Methods. 2021;18(2):203–11. pmid:33288961
  17. 17. Gu A, Dao T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint. 2023.
  18. 18. O R, P F, T B. U-Net: Convolutional networks for biomedical image segmentation. Springer. Cham. 2015. https://doi.org/10.1007/978-3-319-24574-4_28
  19. 19. Kanda Y. Investigation of the freely available easy-to-use software “EZR” for medical statistics. Bone Marrow Transplant. 2013;48(3):452–8. pmid:23208313
  20. 20. Obermeyer Z, Emanuel EJ. Predicting the Future - Big Data, Machine Learning, and Clinical Medicine. N Engl J Med. 2016;375(13):1216–9. pmid:27682033
  21. 21. Mayer E, Herdman G, Burnett C, Kabala J, Goddard P, Potts MJ. Serial STIR magnetic resonance imaging correlates with clinical score of activity in thyroid disease. Eye (Lond). 2001;15(Pt 3):313–8. pmid:11450728
  22. 22. Yoshikawa K, Higashide T, Nakase Y, Inoue T, Inoue Y, Shiga H. Role of rectus muscle enlargement in clinical profile of dysthyroid ophthalmopathy. Jpn J Ophthalmol. 1991;35(2):175–81. pmid:1779487
  23. 23. Bontzos G, Papadaki E, Mazonakis M, Maris TG, Tsakalis NG, Drakonaki EE, et al. Extraocular Muscle Volumetry for Assessment of Thyroid Eye Disease. J Neuroophthalmol. 2022;42(1):e274–80. pmid:34629402
  24. 24. Zhu Y, Wei R, Gao G, Ding L, Zhang X, Wang X, et al. Fully automatic segmentation on prostate MR images based on cascaded fully convolution network. J Magn Reson Imaging. 2019;49(4):1149–56. pmid:30350434
  25. 25. Kurata Y, Nishio M, Kido A, Fujimoto K, Yakami M, Isoda H, et al. Automatic segmentation of the uterus on MRI using a convolutional neural network. Comput Biol Med. 2019;114:103438. pmid:31521902
  26. 26. Haruna Y, Tagami M, Tomita M, Sakai A, Misawa N, Asano K, et al. Correlation Between Changes in Extraocular Muscles and Intraocular Pressure Following Anti-Inflammatory Therapy in Active Thyroid Eye Disease. J Clin Med. 2025;14(5):1480. pmid:40094965