Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

A public generalizable AI tool for automated segmentation of coronal brain tissue slabs for 3D neuropathology

  • Jonathan Williams Ramirez ,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Validation, Visualization, Writing – original draft, Writing – review & editing

    jwilliamsramirez@mgh.harvard.edu

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Dina Zemlyanker,

    Roles Data curation, Formal analysis, Investigation, Methodology, Validation, Writing – original draft

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Lucas Deden-Binder,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Validation

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Rogeny Herisse,

    Roles Data curation

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Erendira Garcia Pallares,

    Roles Data curation, Investigation

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Karthik Gopinath,

    Roles Supervision, Writing – review & editing

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Harshvardhan Gazula,

    Roles Conceptualization, Supervision

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Christopher Mount,

    Roles Data curation, Investigation

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Liana N. Kozanno,

    Roles Data curation, Investigation

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Michael S. Marshall,

    Roles Data curation, Investigation

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Theresa R. Connors,

    Roles Data curation, Investigation

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Matthew P. Frosch,

    Roles Funding acquisition, Resources

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Mark Montine,

    Roles Data curation, Resources

    Affiliation Department of Laboratory Medicine and Pathology, University of Washington School of Medicine, Seattle, Washington, United States of America

  • Derek H. Oakley,

    Roles Funding acquisition, Resources, Supervision

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Christine L. Mac Donald,

    Roles Funding acquisition, Resources

    Affiliation Department of Neurological Surgery, University of Washington School of Medicine, Seattle, Washington, United States of America

  • C. Dirk Keene,

    Roles Funding acquisition, Resources

    Affiliation Department of Laboratory Medicine and Pathology, University of Washington School of Medicine, Seattle, Washington, United States of America

  • Bradley T. Hyman,

    Roles Funding acquisition, Resources, Supervision

    Affiliation Massachusetts Alzheimer’s Disease Research Center, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  • Bruce Fischl,

    Roles Funding acquisition, Supervision, Writing – review & editing

    Affiliation Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America

  •  [ ... ],
  • Juan E. Iglesias

    Roles Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliations Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, Boston, Massachusetts, United States of America, Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Boston, Massachusetts, United States of America, Hawkes Institute, University College London, London, United Kingdom

  • [ view all ]
  • [ view less ]

Abstract

Advances in image registration and machine learning have recently enabled volumetric analysis of postmortem brain tissue from conventional photographs of coronal slabs, which are routinely collected in brain banks and neuropathology laboratories around the world. One caveat of this methodology is the requirement of segmentation of the tissue from the background and out-of-slice tissue in photographs, which currently requires laborious manual intervention. Manual delineation is a bottleneck in this process and poses challenges in scalability, and resources, restricting adoption of these methods. In this article, we present a deep learning model to automate this process. The automatic segmentation tool relies on a U-Net architecture that was trained with a combination of 1,414 manually segmented images of both fixed and fresh tissue, from specimens with varying diagnoses, photographed at two different sites. Automated model predictions on a subset of photographs not seen in training were analyzed to estimate performance compared to manual labels, including both inter- and intra-rater variability. Our model achieved a median Dice score over 0.98, mean surface distance under 0.4 mm, and 95% Hausdorff distance under 1.60 mm, which approaches inter-/intra-rater levels. Our tool is publicly available at surfer.nmr.mgh.harvard.edu/fswiki/PhotoTools and training data is available at https://zenodo.org/records/20647553.

Introduction

Motivation

The fundamental mechanisms underlying many neurodegenerative diseases remain poorly understood, largely due to their molecular complexity, clinical comorbidities, overlapping phenotypes, and pathological heterogeneity [1,2]. As the discipline dedicated to studying and diagnosing diseases of the nervous system through the examination of brain tissue, Neuropathology plays a central role in addressing these challenges.

While certain aspects of neuropathology directly support clinical care, such as guiding therapeutic decisions through diagnosis and prognosis (e.g., tissue biopsies) [35], much of the research aimed at understanding disease pathogenesis, progression, and characterization is conducted postmortem via autopsy [6,7]. Brain autopsy is particularly critical due to the limited accessibility of brain tissue ante mortem and the unique insights gained through histological analysis. Moreover, for many brain diseases, reliable in vivo biomarkers remain lacking, making definitive diagnosis during life difficult or impossible. As a result, neuropathological confirmation at autopsy remains the gold standard for diagnosis [8].

During autopsy, the brain is carefully extracted from the cadaver for subsequent analysis by dissection. Depending on the intended downstream applications, the tissue may be either frozen or chemically fixed (e.g., in formalin) [9]. Dissection typically begins with a gross examination and photographic documentation of the specimen following extraction and fixation, followed by sectioning into slabs. The cerebrum, cerebellum, and brainstem are slabbed separately, and the attending neuropathologist records any macroscopic abnormalities suggestive of underlying pathology [10]. Based on the slab size and features of interest, tissue may then be further subdivided into blocks for histological or molecular processing.

Neuropathologists now have an increasing array of tools for tissue analysis, ranging from conventional histological stains [11] to advanced molecular techniques, including RNA sequencing and multi-omics approaches such as spatial transcriptomics applied to frozen tissue [12]. There is growing interest in integrating these molecular and histopathological data with neuroimaging modalities obtained antemortem (e.g., magnetic resonance imaging, computed tomography, or positron emission tomography) to identify imaging biomarkers that could ultimately enable diagnosis and monitoring in vivo [13,14].

Magnetic resonance imaging (MRI) is widely regarded as the gold standard for morphometric analysis of brain specimens due to its high soft tissue contrast, excellent spatial resolution, and non-invasive nature. However, antemortem MRI is often unavailable or acquired long before death, during which time substantial anatomical changes may occur, limiting its utility for direct comparison with postmortem histopathology. Cadaveric MRI presents a viable alternative, offering the ability to scan intact brains shortly after death. However, unless the brain has been fixed, MRI requires rapid acquisition prior to tissue degradation that could compromise downstream analyses [15]. Such a capability is typically restricted to specialized facilities equipped with dedicated research scanners and appropriate infrastructure. An additional option is ex vivo MRI, which provides greater flexibility in scheduling but necessitates specialized expertise and customized protocols that may not be readily available [16]. Despite these advances, most brain banks and neuropathology laboratories worldwide lack access to MRI scanning resources altogether. While positron emission tomography (PET) enables functional imaging capabilities not possible with MRI, it is restricted to in vivo use and remains limited in availability.

Recently, we proposed a cost-effective alternative to MRI based on dissection photography [13]. Nearly all brain banks and neuropathology laboratories routinely capture photographs of coronal tissue slabs during dissection for documentation purposes [17]. Our method leverages these standard photographs of the cerebral hemispheres to reconstruct 3D volumes, which are then segmented into 11 distinct regions of interest (ROIs) per hemisphere using a neural network. Building on these 3D reconstructions, subsequent work has applied deep learning methods for surface based analysis [18] and tissue slice imputation [19]. This approach provides a practical and accessible alternative to both in situ and ex vivo MRI, significantly reducing cost and acquisition time. Moreover, it enables retrospective analyses of existing datasets, unlocking the potential of legacy collections for new research applications. Even in settings where MRI or PET imaging is available, gross dissection photographs can serve as a steppingstone between the macroscopic detail of in vivo imaging and the microscopic resolution of histology [20,21].

Prior to 3D reconstruction, gross tissue photographs must be preprocessed. First, all photographs must be corrected for pixel size and, ideally, perspective. This ensures that every pixel corresponds to a square of the same size. Next, each slab face must be segmented (often manually) to isolate them from the background and to remove any visible pial surface outside the cut plane. Finally, the slab order and connected components must be defined, critical for disjoint tissue slabs, commonly seen in slabs which include the temporal pole. Steps for image preprocessing are summarized in Fig 1. We have made tools available to facilitate image preprocessing using a graphical user interface created specifically for this task (surfer.nmr.mgh.harvard.edu/fswiki/PhotoTools).

thumbnail
Fig 1. Summary of the image preprocessing pipeline of slab photographs prior to 3D reconstruction.

Images are corrected for pixel size, the slab face is labeled, and the images are masked. Finally, connected components are selected to index slab order and define disjoint tissue slabs. The publicly available tool described in this paper automatizes the Binary segmentation step.

https://doi.org/10.1371/journal.pone.0355740.g001

A major challenge in scaling 3D analysis of dissection photography is thus the need for segmenting brain tissue from the background and out-of-slice tissue, which currently requires labor-intensive manual annotation. This is primarily due to the presence of the cortical surface in the images (Fig 2), which is not part of the slab face and should be excluded from downstream analyses [13]. Segmenting these non-relevant regions is particularly time-consuming, as conventional thresholding techniques (most often effective against high-contrast backgrounds) fail to accurately distinguish the cortical rim from relevant tissue. As a result, manual delineation is often necessary. Automating this segmentation step would significantly reduce the manual workload and enhance the scalability and reproducibility of the pipeline, making high-throughput analysis of dissection photographs more feasible.

thumbnail
Fig 2. Close-up of a dissection photograph (a) and its corresponding reference segmentation (b).

The segmentation excludes the cortical surface (shown in gray), which lies outside the plane of the block face and is not relevant for downstream analysis. Segmenting this region is often time-consuming, as simple thresholding (effective for other parts of the image) fails to distinguish it accurately. Therefore, manual tracing is typically required, making the process labor-intensive. Consequently, fully automated segmentation methods are highly desirable to improve efficiency and scalability.

https://doi.org/10.1371/journal.pone.0355740.g002

Contribution

In this work, we leverage modern deep learning techniques to develop and publicly release an automated segmentation tool for coronal slab photographs. Our method is based on the nnU-Net framework [22], which combines heuristic-based configuration and internal validation to automatically adapt a U-Net architecture [23] to a given segmentation task. nnU-Net has demonstrated state-of-the-art performance across a wide range of 2D and 3D medical imaging challenges [22,24].

We trained our model using 1,414 real photographs from three sources, comprising fresh and fixed tissue. Our experiments show that the model achieves segmentation performance on par with manual annotations. Furthermore, even when manual corrections are necessary, automated predictions substantially reduce the time required by expert labelers, improving overall labeling efficiency. Importantly, the segmentation tool is publicly available and has been integrated into the FreeSurfer software suite [25], facilitating broad adoption within the neuroimaging and neuropathology communities.

In summary, the main contribution of this work are:

  • We share with the research community a segmentation tool of neuropathology brain slabs as part of the FreeSurfer Photo-Tools software suite surfer.nmr.mgh.harvard.edu/fswiki/PhotoTools, facilitating quantitative analysis of post mortem tissue slabs.
  • We present performance evaluation of the tool comparing it with intra- and inter- rater variability of human labelers. Demonstrating that our tool reaches performance comparable of humans with a median Dice score over 0.98, mean surface distance under 0.4 mm, and 95% Hausdorff distance under 1.60 mm.
  • We share a unique dataset of 1,414 manually labeled brain slab images from two different sites, spanning fresh and fixed tissue. Providing useful training or validation data for future segmentation models.

Further related work

Automated segmentation of ROIs from medical images is a vast field, as segmentation is a prerequisite for important downstream analyses (e.g., volumetry or shape analysis), and manual segmentation requires specific expertise, is not reproducible [26,27], and can be highly time consuming – particularly when performed at scale. In the context of neuropathology, classical methods have traditionally relied on handcrafted features and algorithmic approaches to segment tissue from background, delineate cellular boundaries, and classify cell types. Techniques such as thresholding, edge detection, region growing, and morphological operations have been widely used due to their simplicity and interpretability [28]. More advanced methods, including watershed algorithms [29] and active contours (snakes [30]), have enabled finer delineation of complex anatomical features, especially in histological and microscopy images. These approaches often leveraged highly-specific domain knowledge (e.g., staining characteristics, spatial organization, etc.) to enhance segmentation accuracy [31]. While effective in controlled settings, classical methods typically struggle with variability in staining, tissue artifacts, and complex cellular morphology, motivating the transition towards data-driven techniques based on machine learning for more robust and accurate automated analysis.

Deep learning has rapidly advanced computational neuropathology. Convolutional neural networks (CNNs [32,33]) and more recently transformers [34] enabled accurate and scalable solutions to tasks such as tissue segmentation, cell detection, or cell-type classification. In segmentation, architectures like the ubiquitous U-Net [23] enable automatic delineation of tissue compartments or nuclei with high accuracy. Detection models, including region-based CNNs, focus on identifying individual cells or pathological structures, while classification models typically operate on image patches or single-cell representations to assign cell types or disease labels [35]. Hybrid approaches leveraging the global self-attention of transformers and the computational efficiency and strong local feature extraction of CNN methods [36,37]. More recently, foundation models have introduced new levels of generalizability. Promptable segmentation models such as SAM [38], MedSAM [39], and ScribblePrompt [40], can adapt to a wide range of biomedical images with minimal task-specific training. In digital pathology, large-scale pretrained models have recently demonstrated strong generalizability and performance across tasks [41]. Alongside foundation models, emerging diffusion-based segmentation approaches have also proven to be promising for medical image segmentation, leveraging the iterative refinement to generate segmentations with performance competitive to other state-of-the-art methods [42,43].

Human brain slab photography is a niche application where, to the best of our knowledge, dedicated segmentation methods do not exist. However, existing machine learning-based segmentation approaches are highly effective and can be adapted to this specific problem, provided sufficient annotated training data are available. Despite the growing popularity of vision transformers in medical imaging [44,45], recent evidence demonstrates that well-designed U-Net architectures can still match or even surpass the performance of transformer-based models [24]. Designing an optimal U-Net model involves numerous decisions regarding architecture and hyperparameters, which can be challenging to generalize across different datasets and tasks. The widely adopted nnU-Net framework [22] addresses this challenge by combining heuristic rules and internal validation to automate these design choices. Heuristics, based on dataset properties such as image resolution and dimensionality, guide most configuration parameters and have been shown to generalize effectively across diverse biomedical imaging tasks. Internal validation is reserved primarily for selecting the best ensembling and post-processing strategies, as these choices are less predictable from dataset characteristics alone.

Materials and methods

Datasets

Three photographic datasets were obtained from two distinct sites, each using different imaging setups and tissue types (e.g., single hemisphere vs. whole brain, frozen vs. fixed). All specimens were sectioned in the coronal plane in an anterior-to-posterior orientation and subsequently photographed using an overhead camera. Each image contains a variable number of coronal slabs, imaged from the posterior side, and includes a spatial reference object (e.g., rulers, grids, or fiducial markers) to enable pixel size scaling and, when possible, perspective correction.

Massachusetts Alzheimer’s Disease Research Center (MADRC) (733 total photographs)

This dataset includes photographs of fixed tissue slabs from 22 whole brains, 66 left hemispheres, and 41 right hemispheres (65 males, 64 females), collected at the MADRC, affiliated with Massachusetts General Hospital. The cohort primarily consists of individuals diagnosed with neurodegenerative diseases and is skewed toward older donors (average age at death: 7414 years). Over the data collection period, several slabbing protocols were employed, resulting in variability such as differences in slab thickness. Following dissection, slabs were photographed with a ruler placed in the frame to provide a reference for pixel size estimation [13]. An example image from this dataset is shown in Fig 3a.

thumbnail
Fig 3. Example dissection photography setups across the three photographic datasets used in this study.

(a) Photograph from the MADRC dataset showing formalin-fixed slabs from a whole brain, with a ruler for scale calibration. (b) Image from the UW-fixed dataset, also depicting formalin-fixed whole-brain slabs with two orthogonal rulers included for spatial reference. (c,d) Images from the UW-fresh dataset showing coronal slabs from single hemispheres photographed on two slightly different tables. Four fiducial markers placed in a rectangular configuration are used for pixel size and perspective correction.

https://doi.org/10.1371/journal.pone.0355740.g003

University of Washington, Department of Laboratory Medicine and Pathology

Two additional datasets (one using fresh tissue and one using fixed tissue) were acquired from the Department of Laboratory Medicine and Pathology at the University of Washington (UW). For both datasets, coronal slabbing was performed using a modified meat slicer that produced consistent 4 mm slices. Prior to slicing, specimens were embedded in dental alginate to stabilize the tissue and minimize deformation during cutting.

  • UW-fixed (218 total photographs): This dataset consists of dissection photographs from 28 whole brains (17 males, 11 females), fixed in 10% neutral buffered formalin. Similar to the MADRC dataset, many donors in this cohort had neurodegenerative conditions such as Alzheimer’s disease, Lewy body dementia, or Parkinson’s disease (average age at death: 6723 years). Each image includes two orthogonal rulers to enable accurate pixel size calibration. A representative image is shown in Fig 3b.
  • UW-fresh (681 total photographs): This dataset comprises coronal slabs from 38 single hemispheres (19 from males, 19 from females), with donors having an average age at death of 4817 years. This datasets enables us to test our algorithm under more challenging conditions, since fresh tissue (which allows for analyses that are not possible with fixed tissue due to damage to proteins and genetic material, such as RNA sequencing and transcriptomics) suffers from much stronger nonlinear deformation than fixed specimens. Four fiducial markers (on the corners of a rectangle of known dimensions) were used as references for perspective and pixel size estimation [13]. Two slightly different tables (shown in Fig 3c and 3d) were used to acquired this dataset.

Methods

Image preprocessing

Calibrating pixel size in medical image segmentation is essential to ensure that spatial measurements reflect true anatomical dimensions. Eliminating pixel size variability enables better generalization across photographs from different sites, resolutions, or zoom levels. Our preprocessing followed [13] to resample all images to the target resolution (in practice, 0.1 mm). In the MADRC and UW-fixed datasets, we manually clicked on two points along reference rulers and specified the distance between them in physical units (mm). This enabled us to estimate an isotropic scaling factor that was used for resampling. In the UW-fresh dataset, we automatically detected the fiducials with a combination of SIFT [46] and RANSAC [47] algorithms, and used the coordinates to estimate a homography that was used to resample perspective-corrected images at the target resolution.

Manual delineation

Each photograph was manually annotated once by one of four labelers (JWR, LDB, RH, EGP) using open-source photo editing software (GIMP v2.8). Annotators removed all tissue outside the slab surface, including the cortical surface (Fig 2), which is not part of the slab face and should not be considered in subsequent analyses. Therefore, precise masking of the cortical surface in the training data is important to train a model that can accurately segment it at test time – particularly in datasets with thick slabs that expose larger portions of the cortex, as in Fig 2a. Manual segmentation combined thresholding (which efficiently captures the majority of boundaries) with subsequent manual boundary tracing to refine the contours where needed, typically along the cortical surfaces described earlier.

Model training and architecture

We trained our neural network using 733 photographs from the MADRC dataset and 681 from the UW-fresh dataset. The UW-fixed dataset was reserved exclusively for testing, allowing us to evaluate the model’s ability to generalize to previously unseen data. We note that nnU-net internally partitions the training data into training and validation splits, to combat overfitting.

To optimize inference speed while preserving anatomical detail, all images were downsampled to a resolution of 0.5 mm/pixel, striking a balance between spatial resolution and computational efficiency. The average image dimensions used during training were 440 673 pixels.

The final U-Net architecture (automatically configured by the nnU-Net framework) consisted of seven stages with 32, 64, 128, 256, 512, 512, and 512 feature channels, respectively. Model training was performed over five cross-validation folds, each trained for 1,000 epochs, and required a total of 58.3 hours on an NVIDIA RTX A6000 GPU.

Evaluation metrics

To evaluate automated segmentations with respect to the reference masks, we use modified versions of three standard metrics: Dice overlap, average symmetric surface distance, and the 95% Hausdorff Distance (HD95). The modification lies in the fact that we only compute the metric over pixels within 10 mm of the (morphologically filled) reference segmentations. This is to reflect the fact that we do not mind false positives far away from the target slabs (e.g., due to the presence of non-target slabs, as in Fig 5f or S3 and S4 Figs in the supplement), and which can be easily masked during postprocessing. In Fig 4, we show how users can select connected components using the graphical user interface on our publicly available tools, easily omitting any false positives in model predictions.

thumbnail
Fig 4. Example snapshot of our graphical user interfaces selecting only the relevant connected components, easily removing any false positives from the predicted label.

Left: Original image masked using the model prediction. Right: Manually selected connected components, omitting false positives from the model prediction.

https://doi.org/10.1371/journal.pone.0355740.g004

thumbnail
Fig 5. (a) Sample photograph from the MADRC dataset.

(b) Corresponding reference mask. (c) Automated segmentation. (d-f) Example from UW-fixed dataset; note the non-target slabs leading to irrelevant false positives that are factored out of our accuracy metrics. (g-h) Example from the UW-fresh dataset. A more comprehensive set of examples can be found in the supplement S1 to S6 Figs.

https://doi.org/10.1371/journal.pone.0355740.g005

Ethics statement

Human tissue used in this study was obtained in accordance with institutional and national ethical guidelines. Procedures for tissue collection and processing at Massachusetts General Hospital were reviewed and approved by the Partners Human Research Committee institutional review board under protocol 1999P009556. All postmortem specimens were obtained with informed consent from next of kin. At the University of Washington, tissue processing and collection involves the receipt and analysis of de-identified information and specimens from deceased individuals which were acquired from a biorepository. Consent for the use of information and specimens from the deceased individuals was obtained either from the individual while they are alive, or from the individual’s legally authorized representative. Procurement and banking practices of the biorepository are informed by the US Revised Uniform Anatomical Gift ACT 2006 (Last Revised or Amended in 2009) and Washington Statute Chapter 68.64 RCW. We work closely with the UW School of Medicine Compliance Office on our consent forms and HIPAA compliance. All materials are collected under informed consent. All data was accessed for research purposes from 07/08/2022–09/05/2025.

Experiments and results

Manual segmentation: Time efficiency and inter-/intra-rater variability

To contextualize the performance of the automated method, we first conducted an experiment involving only human labelers, with two main objectives: to assess labeling time efficiency (i.e., the time required for manual annotation), and to evaluate variability both within and between annotators (i.e., intra- and inter-rater variability).

Labeling time was measured on a subset of 20 photographs. On average, manual annotation took 3 minutes and 29 seconds per tissue slab, implying that a typical photograph containing three slabs required approximately 10 minutes to label. Extrapolating this estimate to the entire training dataset of 1,414 images yields a total of 214.9 hours of continuous manual labeling. These results underscore the potential of the automated approach to substantially reduce expert annotation time and associated resource costs.

To assess annotation variability, 20 photographs (10 from the MADRC dataset and 10 from the UW-fresh dataset) were randomly selected and annotated twice by two labelers (JWR and LDB), with a one-week interval between sessions to minimize memory bias. Fig 6 presents the Dice scores, average symmetric surface distances, and HD95 for both intra- and inter-rater comparisons. For inter-rater evaluation, results were averaged across all four possible labeler pairings. All metrics indicated excellent agreement, with median Dice scores exceeding 0.98, average surface distances below 0.5 mm, and HD95 values under 2 mm. Notably, the inter-rater variability was comparable to the intra-rater variability of the more variable annotator, demonstrating strong overall consistency and reliability in the manual annotations.

thumbnail
Fig 6. Box plots for Dice scores, mean surface distance, and HD95 for manual and automated methods.

“Intra-1”: intra-rater variability of Labeler 1, measured on 20 images. “Intra-2”: intra-rater variability of Labeler 2 on the same 20 images. “Inter”: inter-rater variability of the same 20 images. “Auto-in”: accuracy of the proposed automated method on the in-distribution data (159 photographs). “Auto-out”: accuracy of the proposed automated method on the out-of-distribution data (218 photographs).

https://doi.org/10.1371/journal.pone.0355740.g006

Automated segmentation: In- and out-of-distribution images

We computed Dice scores and surface distance metrics for two evaluation settings: (i) 89 photographs from the MADRC dataset and 70 from the UW-fresh dataset that were withheld during training, and (ii) the full UW-fixed dataset, comprising 218 photographs. The first setting evaluates model performance on in-distribution data, while the second assesses generalization to out-of-distribution images.

Some segmentations are shown in Fig 5 and in the supplement S1 to S6 Figs, whereas quantitative results are shown in Fig 6. On in-distribution data, the model achieved excellent performance, with a median Dice score of 0.985, a mean surface distance of 0.34 mm, and a HD95 of 1.27 mm. These values closely approach the intra-rater variability of the more consistent human annotator (noting that this comparison involves different image samples). Using a very conservative threshold of 2 mm for the mean surface distance to define outliers, the in-distribution outlier rate was just 2%.

On the out-of-distribution UW-fixed dataset, the model maintained strong performance, with a median Dice score of 0.981, a mean surface distance of 0.37 mm, and an HD95 of 1.60 mm, closely matching results on in-distribution data. These results indicate that the model generalizes well to previously unseen data in the majority of cases. As expected with domain shifts, variability increased slightly: approximately 6% of images exhibited errors greater than 2 mm, as seen in the wider spread of the box plots across all metrics. However, this increase is modest in absolute terms, particularly given the relatively stringent threshold used to define outliers. Overall, the model demonstrates robust generalization with only minimal degradation in performance on out-of-distribution samples. To determine modes of failure, we examined the worst performing images in our out-of-distribution data. Among the worst performing images, the main driver for error came from ambiguous regions that resulted in high variability in manual annotations. This occurred most frequently along the pial surface (Fig 7), where the resulting discrepancies were often minor and could be corrected efficiently when very high accuracy was required.

thumbnail
Fig 7. Example image for worst performing out-of-distribution fixed tissue labels.

Left: Original dissection photograph. Right: Overlaid segmentations showing true positives (green), false negatives (red), and false positives (orange). We note that the false positives can be cleaned up manually with very few clicks.

https://doi.org/10.1371/journal.pone.0355740.g007

Discussion

This work introduces a fully automated segmentation tool for brain dissection photographs, offering a practical and scalable alternative to resource-intensive annotation workflows. By leveraging the nnU-Net framework, our method achieves a level of accuracy that is comparable to that of expert human annotators. It yields strong performance on in-distribution data, and generalizes effectively to out-of-distribution images from a previously unseen site.

Despite these encouraging results, a key limitation of this study is the diversity of the data sources. All training and evaluation were conducted on photographs acquired from only three sites. While the diversity in dissection protocols, imaging setups, and tissue conditions (e.g., fixed vs. fresh) helps build robustness, it is still limited in scope. For best performance we recommend for users to use a high-contrast black background, with an overhead camera angle perpendicular to the dissection table. Because of variability lighting conditions in the training data, we expect that the tool will be more resistant to lighting changes in out-of-distribution data compared to changes that more dramatically change the statistical properties of the images, such as nonstandard backgrounds. Further validation on additional datasets (particularly from institutions with different photographic and tissue-handling practices) will be important to assess and potentially improve the generalizability of the tool.

To this end, future work may explore additional strategies to enhance robustness across domains. For example, implementing photographic data from more sites with different photographic setups will improve overall generalizability. Additionally, training a hybrid model including synthetic data generation with domain randomization could further expose the model to a broader range of appearances and imaging conditions. In the synthetic context we could model for variables like perspective, lighting, as well as tissue deformation. Incorporating style transfer or domain adaptation techniques may also reduce sensitivity to site-specific artifacts. Another potential direction is to use lightweight image-level quality control (QC) mechanisms, such as binary pass/fail ratings, to complement traditional spatial metrics such as Dice and surface distance. Such QC frameworks may provide more intuitive assessments of segmentation reliability in large-scale, real-world workflows. Future work will explore how other parts of the preprocessing process can be automated. For example, the current need for manual correction of retrospective data with be circumvented via automatic detection of reference landmarks (e.g., rulers), while keeping the opportunity for manual editing when needed.

Overall, our findings demonstrate that automated segmentation can replace a time-intensive manual step while maintaining expert-level accuracy across diverse brain dissection photographs. By making high-throughput analysis of slab photograph archives more feasible, this tool may help unlock large retrospective neuropathology datasets that would otherwise remain difficult to quantify at scale: the practical impact of this tool ultimately lies in its potential to dramatically reduce the time burden of manual annotation. As we estimated, segmenting a typical dissection photograph takes approximately 10 minutes by hand. Extrapolated to our training dataset alone, this represents over 200 hours of expert time. Our tool can almost entirely remove this workload for most images and still substantially accelerate expert review when manual refinement is needed. Additionally, sharing of manual ground truth labels will encourage others in the community to experiment with other segmentation methods and share their data. This represents a significant step toward scalable, high-throughput analysis of neuropathology specimens, unlocking the potential of vast image archives that would otherwise be prohibitively time-consuming to label manually.

Supporting information

S1 Fig. Additional automated segmentations for sample images from MADRC dataset (in distribution).

https://doi.org/10.1371/journal.pone.0355740.s001

(TIFF)

S2 Fig. Additional automated segmentations for sample images from MADRC dataset (in distribution).

https://doi.org/10.1371/journal.pone.0355740.s002

(TIFF)

S3 Fig. Additional automated segmentations for sample images from UW-fixed dataset (out of distribution).

Please note that we do not mind false positives on non-target slabs, which can be easily masked during postprocessing.

https://doi.org/10.1371/journal.pone.0355740.s003

(TIFF)

S4 Fig. Additional automated segmentations for sample images from UW-fixed dataset (out of distribution).

Please note that we do not mind false positives on non-target slabs, which can be easily masked during postprocessing.

https://doi.org/10.1371/journal.pone.0355740.s004

(TIFF)

S5 Fig. Additional automated segmentations for sample images from UW-fresh dataset (in distribution).

https://doi.org/10.1371/journal.pone.0355740.s005

(TIFF)

S6 Fig. Additional automated segmentations for sample images from UW-fresh dataset (in distribution).

https://doi.org/10.1371/journal.pone.0355740.s006

(TIFF)

Acknowledgments

The authors would like to thank the research participants and their families without whom this work would be impossible.

References

  1. 1. Jellinger KA. Basic mechanisms of neurodegeneration: a critical update. J Cell Mol Med. 2010;14(3):457–87. pmid:20070435
  2. 2. Armstrong RA, Lantos PL, Cairns NJ. Overlap between neurodegenerative disorders. Neuropathology. 2005;25(2):111–24. pmid:15875904
  3. 3. Ahn ES, Goumnerova L. Endoscopic biopsy of brain tumors in children: diagnostic success and utility in guiding treatment strategies. J Neurosurg Pediatr. 2010;5(3):255–62. pmid:20192642
  4. 4. Vital A, Vital C. Clinical Neuropathology practice guide 3-2014: combined nerve and muscle biopsy in the diagnostic workup of neuropathy - the Bordeaux experience. Clin Neuropathol. 2014;33(3):172–8. pmid:24618073
  5. 5. Weis J, Brandner S, Lammens M, Sommer C, Vallat J-M. Processing of nerve biopsies: a practical guide for neuropathologists. Clin Neuropathol. 2012;31(1):7–23. pmid:22192700
  6. 6. Sejda A, Wierzba-Bobrowicz T, Adamek D, Gulczyński J, Michalak S, Grajkowska W, et al. Central nervous system autopsy - a neuropathological procedure based on multidisciplinary pathoclinical cooperation. Neurol Neurochir Pol. 2022;56(2):118–30. pmid:34913473
  7. 7. Oura P, Hakkarainen A, Sajantila A. Forensic neuropathology in the past decade: a scoping literature review. Forensic Sci Med Pathol. 2024;20(2):724–35. pmid:37439948
  8. 8. Hilton DA, Shivane AG. Neurodegenerative diseases. Springer Nature Switzerland. Cham, Switzerland. 2015.
  9. 9. Adams JH, Murray MF. The brain. Cambridge University Press; 1982.
  10. 10. Adams JH, Murray MF. Dissection of the fixed brain. Cambridge University Press; 1982.
  11. 11. Garman RH. Histology of the central nervous system. Toxicol Pathol. 2011;39(1):22–35. pmid:21119051
  12. 12. Grima N, Henden L, Watson O, Blair IP, Williams KL. Simultaneous isolation of high-quality RNA and DNA from postmortem human central nervous system tissues for omics studies. J Neuropathol Exp Neurol. 2022;81(2):135–45. pmid:34939123
  13. 13. Gazula H, Tregidgo HFJ, Billot B, Balbastre Y, Williams-Ramirez J, Herisse R, et al. Machine learning of dissection photographs and surface scanning for quantitative 3D neuropathology. Elife. 2024;12:RP91398. pmid:38896568
  14. 14. Tregidgo HFJ, Casamitjana A, Latimer CS, Kilgore MD, Robinson E, Blackburn E, et al. 3D reconstruction and segmentation of dissection photographs for MRI-free neuropathology. Lecture Notes in Computer Science. Springer International Publishing; 2020. p. 204–14. https://doi.org/10.1007/978-3-030-59722-1_20
  15. 15. Guo Y, Ma J, Ma X, Ye K, Jiang C, Xu J. Single-nucleus RNA sequencing revealed the impact of post-mortem interval on the cellular component and gene expression analysis of mouse brains. bioRxiv. 2025;2024–12.
  16. 16. Edlow BL, Mareyam A, Horn A, Polimeni JR, Witzel T, Tisdall MD, et al. 7 Tesla MRI of the ex vivo human brain at 100 micron resolution. Sci Data. 2019;6(1):244. pmid:31666530
  17. 17. Rampy BA, Glassy EF. Pathology gross photography: the beginning of digital pathology. Surg Pathol Clin. 2015;8(2):195–211. pmid:26065794
  18. 18. Gopinath K, Greve D, Das S, Arnold S, Magdamo C, Iglesias JE. Cortical analysis of heterogeneous clinical brain MRI scans for large-scale neuroimaging studies. International Conference on Medical Image Computing and Computer-Assisted Intervention 2023:35–45.
  19. 19. Aguirre M, Williams-Ramirez J, Zemlyanker D, Hu X, Deden-Binder L, Herisse R. Improving neuropathological reconstruction fidelity via AI slice imputation. 2026. https://arxiv.org/abs/2602.00669
  20. 20. Pichat J, Iglesias JE, Yousry T, Ourselin S, Modat M. A survey of methods for 3D histology reconstruction. Med Image Anal. 2018;46:73–105. pmid:29502034
  21. 21. Athalye C, Bahena A, Khandelwal P, Emrani S, Trotman W, Levorse LM, et al. Operationalizing postmortem pathology-MRI association studies in Alzheimer’s disease and related disorders with MRI-guided histology sampling. Acta Neuropathol Commun. 2025;13(1):120. pmid:40437594
  22. 22. Isensee F, Jaeger PF, Kohl SAA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat Methods. 2021;18(2):203–11. pmid:33288961
  23. 23. Ronneberger O, Fischer P, Brox T. U-net: convolutional networks for biomedical image segmentation. Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer; 2015. p. 234–41.
  24. 24. Isensee F, Wald T, Ulrich C, Baumgartner M, Roy S, Maier-Hein K, et al. nnu-net revisited: a call for rigorous validation in 3d medical image segmentation. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer; 2024. p. 488–98.
  25. 25. Fischl B. FreeSurfer. Neuroimage. 2012;62(2):774–81.
  26. 26. Pham DL, Xu C, Prince JL. Current methods in medical image segmentation. Annu Rev Biomed Eng. 2000;2:315–37. pmid:11701515
  27. 27. Liu X, Song L, Liu S, Zhang Y. A review of deep-learning-based medical image segmentation methods. Sustainability. 2021;13(3):1224.
  28. 28. Gonzalez RC. Digital image processing. Pearson; 2009.
  29. 29. Vincent L, Soille P. Watersheds in digital spaces: an efficient algorithm based on immersion simulations. IEEE Trans Pattern Anal Machine Intell. 1991;13(6):583–98.
  30. 30. Kass M, Witkin A, Terzopoulos D. Snakes: active contour models. International journal of computer vision. 1988;1(4):321–31.
  31. 31. Meijering E. Cell segmentation: 50 years down the road. IEEE Signal Processing Magazine. 2012;29(5):140–5.
  32. 32. LeCun Y, Boser B, Denker JS, Henderson D, Howard RE, Hubbard W, et al. Backpropagation applied to handwritten zip code recognition. Neural Computation. 1989;1(4):541–51.
  33. 33. Lecun Y, Bottou L, Bengio Y, Haffner P. Gradient-based learning applied to document recognition. Proc IEEE. 1998;86(11):2278–324.
  34. 34. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN. Attention is all you need. Advances in neural information processing systems. 2017;30.
  35. 35. Gamper J, Alemi Koohbanani N, Benet K, Khuram A, Rajpoot N. PanNuke: an open pan-cancer histology dataset for nuclei instance segmentation and classification. Lecture Notes in Computer Science. Springer International Publishing; 2019. p. 11–9. https://doi.org/10.1007/978-3-030-23937-4_2
  36. 36. Chen J, Lu Y, Yu Q, Luo X, Adeli E, Wang Y. Transunet: transformers make strong encoders for medical image segmentation. 2021. https://arxiv.org/abs/2102.04306
  37. 37. Li J, Chen J, jiang L, Li R, Han P, Cheng J. SEAformer: selective Edge Aggregation transformer for 2D medical image segmentation. Biomedical Signal Processing and Control. 2025;102:107203.
  38. 38. Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L. Segment anything. arXiv preprint. 2023. https://doi.org/arXiv:230402643
  39. 39. Ma J, He Y, Li F, Han L, You C, Wang B. Nature Communications. 2024;15(1):654.
  40. 40. Wong HE, Rakic M, Guttag J, Dalca AV. ScribblePrompt: Fast and Flexible Interactive Segmentation for Any Biomedical Image. European Conference on Computer Vision (ECCV 2024). Berlin, Heidelberg: Springer; 2024. p. 207–29.
  41. 41. Xu H, Usuyama N, Bagga J, Zhang S, Rao R, Naumann T, et al. A whole-slide foundation model for digital pathology from real-world data. Nature. 2024;630(8015):181–8. pmid:38778098
  42. 42. Wu J, Ji W, Fu H, Xu M, Jin Y, Xu Y. MedSegDiff-V2: diffusion-based medical image segmentation with transformer. AAAI. 2024;38(6):6030–8.
  43. 43. Xing Z, Wan L, Fu H, Yang G, Yang Y, Yu L, et al. Diff-UNet: a diffusion embedded network for robust 3D medical image segmentation. Med Image Anal. 2025;105:103654. pmid:40602205
  44. 44. Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, et al. A survey on vision transformer. IEEE Trans Pattern Anal Mach Intell. 2023;45(1):87–110. pmid:35180075
  45. 45. Khan S, Naseer M, Hayat M, Zamir SW, Khan FS, Shah M. Transformers in vision: a survey. ACM Comput Surv. 2022;54(10s):1–41.
  46. 46. Lowe DG. Object recognition from local scale-invariant features. In: Proceedings of the Seventh IEEE International Conference on Computer Vision. 1999. p. 1150–7 vol. 2. https://doi.org/10.1109/iccv.1999.790410
  47. 47. Fischler MA, Bolles RC. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM. 1981;24(6):381–95.