Fig 1.
Comparison of the Swin Transformer and the Vision Transformer.
(A) Swin Transformer; (B)Vision Transformer.
Fig 2.
Swin Transformer structure.
Fig 3.
Two consecutive swin transformer blocks.
Fig 4.
Self-attention computed based on shifted windows.
Fig 5.
Calculation of self-attention of moving windows based on cyclic shifting.
Fig 6.
Abstract view of object detection systems.
Fig 7.
SwinT-YOLOv4 structure.
Fig 8.
SPP network structure.
Fig 9.
PANet network structure.
Table 1.
Setting of training parameters.
Fig 10.
SwinT-YOLOv4 loss function curve.
Table 2.
Comparison of YOLOv4 and SwinT-YOLOv4 results.
Fig 11.
mAP of SwinT-YOLOv4 and YOLOv4.
(A) mAP of YOLOv4; (B) mAP of SwinT-YOLOv4.
Table 3.
Comparison of Fps and FLOPs between YOLOv4 and SwinT-YOLOv4.
Fig 12.
Comparison of the detection effect of SwinT-YOLOv4 and YOLOv4.
(A) Detection effect of YOLOv4; (B) Detection effect of SwinT-YOLOv4.
Table 4.
The performance comparison of existing model.