Fig 1.
Schematic of the anatomical structure of a breast ultrasound (BUS) image.
Fig 2.
Illustration of the approaches used for performing BUS image segmentation and classification jointly.
(1) Cascade approach; (2) Multi-task learning approach; (3) Cross-task approach.
Fig 3.
Sample BUS images from the THH, UDIAT, and BUSI datasets with average image size.
Fig 4.
Illustration of the cross-task approach for achieving mutual improvement between tasks.
Fig 5.
Architecture of the cross-task guided network (CTG-Net).
The network comprises a shared features backbone and three units for BUS image classification and segmentation. ASPP, Atrous Spatial Pyramid Pooling; GAP, Global Average Pooling; FC, Fully Connected Layer; Concat, Features Channel Concatenation.
Fig 6.
Schematic of Lesion Attention Module (LAM).
⊙ denotes element-wise multiplication and ⊕ denotes element-wise addition.
Fig 7.
Schematic of the Category Selection Module (CSM).
⊙ denotes element-wise multiplication and ⊗ denotes matrix multiplication. Conv indicates that the features have passed through a convolutional layer.
Fig 8.
Schematic of the Anatomical Knowledge Guidance Module (AKGM).
⊙ denotes element-wise multiplication, ⊗ denotes matrix multiplication and ⊕ denotes element-wise addition. Conv indicates that the features have passed through a convolutional layer.
Fig 9.
Visualization of segmentation results and class activation maps for the BUS image private dataset.
(a) BUS images; (b) ground truth (white areas are lesions); (c)–(g) segmentation results are achieved, by U-Net, Attention U-Net, Nested U-Net, MA U-Net, CTG-Net (ours), respectively; (h)–(l) class activation maps are achieved by AlexNet, VGG16, ResNet18, DenseNet121, and CTG-Net (ours), respectively.
Table 1.
Segmentation and classification performance compared with other state-of-art segmentation models on the THH dataset.
Table 2.
Segmentation performance compared with state-of-the-art segmentation models on the UDIAT and BUSI datasets.
Table 3.
Classification performance compared with state-of-the-art classification models on the UDIAT and BUSI datasets.
Fig 10.
Examples of results compared with the state-of-the-art methods on BUS image public datasets.
The state-of-the-art results are from [54] and the results for CTG-Net are from this present study. (a) BUS images, (b) ground truth (white areas are lesions and text are true class labels). (c)–(m) Segmentation results achieved by (c) FCN-AlexNet, (d) SegNet, (e) U-Net, (f) CE-Net, (g) MultiResUNet, (h) RDAU-Net, (i) SCAN, (j) DenseU-Net, (k) STAN, (l) ESTAN, and (m) our CTG-Net.
Table 4.
Segmentation and classification performance compared with state-of-the-art multi-task learning methods on the THH dataset.
Table 5.
Quantitative results for CTG-Net and baseline (i.e., CTG-Net without any task-specific feature extraction module) on the THH dataset.
Table 6.
Ablation study of fine segmentation unit in CTG-Net.
Table 7.
Ablation study of classification unit in CTG-Net.
Fig 11.
Illustration of segmentation failure examples.
(a) BUS Images, (b) ground truth, (c) CTG-Net predictions. Image (1) is from the THH dataset, images (2) and (3) are from the UDIAT dataset, and image (4) is from the BUSI dataset. The text represents the category labels.