Figures
Abstract
Cross-domain object detection is a key problem in the research of intelligent detection models. Different from lots of improved algorithms based on two-stage detection models, we try another way. A simple and efficient one-stage model is introduced in this paper, comprehensively considering the inference efficiency and detection precision, and expanding the scope of undertaking cross-domain object detection problems. We name this gradient reverse layer-based model YOLO-G, which greatly improves the object detection precision in cross-domain scenarios. Specifically, we add a feature alignment branch following the backbone, where the gradient reverse layer and a classifier are attached. With only a small increase in computational, the performance is higher enhanced. Experiments such as Cityscapes→Foggy Cityscapes, SIM10k→Cityscape, PASCAL VOC→Clipart, and so on, indicate that compared with most state-of-the-art (SOTA) algorithms, the proposed model achieves much better mean Average Precision (mAP). Furthermore, ablation experiments were also performed on 4 components to confirm the reliability of the model. The project is available at https://github.com/airy975924806/yolo-G.
Citation: Wei J, Wang Q, Zhao Z (2023) YOLO-G: Improved YOLO for cross-domain object detection. PLoS ONE 18(9): e0291241. https://doi.org/10.1371/journal.pone.0291241
Editor: Praveen Kumar Donta, TU Wien: Technische Universitat Wien, AUSTRIA
Received: June 8, 2023; Accepted: August 24, 2023; Published: September 11, 2023
Copyright: © 2023 Wei et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The project code has been uploaded to Github (https://github.com/airy975924806/yolo-G). The data used in the article has been uploaded to Kaggle (https://www.kaggle.com/datasets/airy975924806/city-foggycity-yoloformat).
Funding: The authors received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Deep convolutional models significantly improve the precision of object detection [1–11]. However, the models are severely constrained by training data. Cross-domain object detection requires the model to fulfill the training process in a fully annotation training set and then applied in the validation from a different domain. Poor results or even degradation is likely to come following due to such a domain gap. Well, in the real world, such differences, roughly lighting conditions, weather conditions, view angles, equipment differences, etc. [12], are quite common. Furthermore, it is impossible to provide infinite training data. That is to say, gathering more data is not a reliable way to improve the ability and enhance the robustness of the model.
Addressing the cross-domain object detection problem within limited labeled training data, DAF [13] pays attention to this problem for the first time, and an improved two-stage detection model based on Faster R-CNN [10] is designed, which took an unsupervised training method to upgrade its detection ability under changing weather conditions. Following up, in [14–23] et al., the addition of guidance branches, fusion generative adversarial model, and other methods are used to gain a much better ability. In [16, 18, 24, 25], semi-supervised methods are adopted, such as adjusting the training strategy, adding the knowledge distillation mechanism [26], or introducing an iterative training way, these all intend to gradually improve the model cross-domain detection ability. Well, these methods significantly increase the budget of computation and lengthen the training time.
Compared with the two-stage detection model, YOLO has more advantages in terms of speed, precision, and application. At present, YOLO-based cross-domain detection models are developing rapidly [27–33]. Focusing on the cross-domain object detection problem, this paper takes the YOLOv5-L model as the baseline. Especially, we adopt the idea of feature alignment which is widely used in two-stage models, and add unsupervised adversarial training branches to improve the adaptive ability of the model. By adopting such a simple and efficient branch, the model gains fresh new abilities. Specifically, to realize the adaptive feature alignment and reduce the domain gap between different images [12, 13], we introduce the adversarial training mechanism into the model parameter update process. Especially, to release the burden, we use a naive three-layer classifier based on the full convolution network and global average maximum pooling, with gradient reverse [34, 35] operation. Compared with the dense linear layer, the YOLO-G just slightly increases the amount of computation, however, the precision is greatly improved. To verify the proposed model in cross-domain object detection tasks, this paper carries out extensive experiments under 6 benchmark sets. The results show that YOLO-G achieves better detection precision than a series of semi-supervised and two-stage SOTA models. The main contributions of this paper are as follows:
- For cross-domain object detection tasks, we verify the usability of the YOLO model in cross-domain object detection tasks through comprehensive experiments. Our ablation experiments show that under the source-only condition, the YOLOV5-L model can compare with many SOTA algorithms.
- The YOLO-G model is designed based on YOLOV5-L. A concise and efficient unsupervised adversarial training branch is added to the baseline model. In this way, we expand the scope of cross-domain model design and change the previous situation limited to two-stage models.
- We carry out qualitative and quantitative experiments and compare 10 algorithms in 6 cross-domain benchmarks. In order to fully illustrate the credibility of the model, we also carry out ablation experiments in 2 aspects and 4 factors. Experimental results show that the proposed model has indeed achieved better detection results.
Related work
Object detection
As an important content of computer vision research, object detection models have developed rapidly with the support of deep learning technology in recent years. Among them, the one-stage detection models are represented by YOLO [5, 7, 27] and SSD [8]. Especially, the YOLO model has gradually developed into a rich series with fast infer speed, accurate precision, and simple deployment. The two-stage detection model, represented by Mask R-CNN [9], and Fast R-CNN [10], occupies an important position in academic research because the models adopt the process of separating ROI region generation and classification discrimination, making the model have higher plasticity. Thanks to the finer ROI search strategy, the two-stage models have higher detection precision under the same conditions. At present, with the application of the transformer, the object detection model based on the transformer [11, 36] with larger parameter scales is shining brightly and shows good development prospects.
In actual deployment scenarios, the YOLO series model has better compatibility. Therefore, considering the task requirements, application scenarios, and inference speed, this paper selects the YOLOv5-L model as the baseline and improves it to adapt to cross-domain object detection tasks.
Cross-domain
DAF [13] creatively uses the two-stage detection model to solve the cross-domain object detection problem and proposes a processing method based on global feature alignment and target feature alignment. Inspired by it, subsequent research [14–18, 24, 37] is mostly based on the two-stage detection model. Namely a few, SWDA [38] proposes to use enhanced local features and global features as auxiliary feature alignment, GPA [39] designs a detection framework based on class feature alignment, MAF [40] is different from the feature alignment method, ATF [14] and PA-ATF [41] try to use multi-branch supervised training to improve the cross-domain object detection ability. With the deepening of related research content, there have been attempts based on one-stage detection models, such as EPMDA [28] integrates FCOS [42] modules into YOLO backbone to improve its ability to extract object features in cross-domain images, SSDA-YOLO [12] adds CUT [43] and knowledge distillation mechanism to YOLO-based cross-domain object detection model [32, 44], are also based on YOLO. Inspired by them, this paper also takes the one-stage detection model as the baseline to explore cross-domain object detection tasks. In fact, through the experimental research of this paper, it is shown that with the help of the functional branch with a limited amount of computation, YOLO-G improves the ability of YOLO in the cross-domain object detection task. Furthermore, in some tasks, the result is far higher than in some two-stage detection models.
Cross-domain detection
Cross-domain detection tasks evolve into multiple paths in the development process, such as fully supervised, semi-supervised, and unsupervised methods. Among them, full supervision refers to the construction of full annotation data of the target domain. Through collection and annotation for data augmentation, this method achieves the purpose of improving the generalization. However, this will bring huge labor. In real application scenarios, the data of all possible scenarios cannot be collected, so this method is not advisable. The semi-supervised [15, 16, 45, 46] method proposes to use partially labeled data to guide the training process. They mainly divide the training process into multiple stages. Firstly, an initial model with fully labeled source domain data is trained, and then pseudo-labels on the target domain data are created by the pre-trained model. After setting the confidence threshold, iteratively add the detected sufficiently prepared target domain data into the training set. The gradual guidance improves the generalization ability, but this method is limited by the initial model detection ability. In many cases, the initial model is difficult to provide enough effective pseudo-labels, making it difficult to continue the iterative process. Unsupervised [13, 34, 47–50] uses the powerful self-learning ability of the deep learning model itself, by setting simple boundary conditions or even providing unlabeled data belonging to the target domain, the model can independently learn to identify the key features required of the object, which is more concise and efficient in term of implementation.
Inspired by these principles, based on the YOLOV5-L model, this paper proposes a simple and accurate cross-domain object detection model by adding unsupervised feature alignment function branches.
Method
Firstly, we briefly introduce the YOLOV5 model, which is mainly composed of 3 parts, respectively. The backbone is responsible for feature extraction, composed of C3, CSP, and SPPF, through a variety of series and parallel forms of residual structure, to extract the feature of the input images. The neck mainly completes feature processing and fusion, using FPN [51], and CSP to achieve bottom-up and top-down feature fusion. The head module, divided into 3 layers, completes the task of detecting targets on small, medium, and large scales respectively. Compared with the two-stage model Faster RCNN, which first provides ROI by RPN and then performs classification detection, the YOLO model directly detects ROI regions on the feature map, which is more prominent in processing efficiency.
As shown in Fig 1, YOLO-G adds functional branches after the backbone. Adversarial training is adapted to constrain the target images from different scenarios to achieve feature alignment in the backbone output layer. By doing so, eliminates the problem of feature inconsistency caused by scene, weather, viewing angle, or other factors, and improves the cross-domain detection precision of the object finally.
The main improvement is a simple and efficient branch is attached behind the backbone; the main modules are illustrated in the bottom-left.
Set S represent the source domain, T is the target domain, XS is one dataset collected from the source domain, xs(c,b)∈XS is a single image, c is the category, b is the box location, which contains four scalers. xt∈XT is a sample identified from the target domain, without category and location annotations. The cross-domain object detection task is to learn a detection model f from the source domain, and then directly use it for inference, which expressed as .
The domain gap between different datasets is the inner cause of the cross-domain problem. The H-divergence [52] is used to quantitatively evaluate this difference, which is essential to measure the discriminant error of different data in the same classifier, and the formula is as follows:
(1)
where dH means the calculated H-divergence, h is the shared classifier. In this paper, the classifier is composed of 2 convolutional layers and 1 global average pooling layer. During the training process, the label of the source domain and target domain is set to 0, 1, separately. The classifier only discriminates the domain category of the feature vector, that is h(xm)→{0,1}. xm is the feature vector extracted by the backbone. errs, errt represent the case of misclassification of the classifier. Therefore, the maximum value of misclassification is the upper bound of the difference datasets, when the classifier reaches the optimal classification ability.
In short, the main work is how to reduce the H-divergence between datasets. According to Eq (1), xm is the key factor affecting H-divergence. That means the larger the H-divergence, the bigger difference for the feature, and vice versa. Therefore, this explains the model’s inability to extract consistent features, resulting in significant performance degradation in cross-domain object detection.
To reduce H-divergence, data-level alignment can be adopted, such as data augmentation technology [53, 54], through copy, paste, GAN generation, and other means to generate fake target domain images to enrich the training set. The second is pixel-level alignment [55], such as image style transfer, in cross-weather, lighting scenes. A variety of style transfer models convert the source domain to the target domain, reducing the difference in data distribution. The third is feature-level alignment [56–58], from the perspective of data distribution, feature is a low-dimensional description of the true distribution of the real data. Therefore, this paper mainly explores the method of realizing cross-domain image alignment at the feature level.
Gradient reverse [35] is an unsupervised training method first proposed in the cross-domain scene classification task, and its basic idea is to maximum-minimize the loss of the classifier. That is, to confuse the classification ability of the classifier. with the purpose to achieve the minimize dH. By doing so, the classifier cannot correctly classify the input features. The process is described as follows:
(2)
where fb is the feature extraction module of the object detection model, and it is the backbone of our YOLO-G.
Gradient reverse sets the opposite weights of the feature vectors during forward propagation and backward propagation. The classifier is guided to reduce the classification error when the classifier is trained forward, and the opposite weight factor is applied when the parameters are updated in the backward propagation. In an interactive way, we guide the backbone to update the heading for maximizing the classification error of the classifier.
During the training progress, the object detection model and domain classification model are updated synchronously. The object detection loss function Ldet consists of three parts: object detection loss Lobj, classify loss Lclass, and box regression loss Ldc.
The loss function of the feature-aligned branch is the binary classify loss Ldc using BCE loss:
(3)
where dk is the domain label of the sample with k index in the training batch, pk(i,j) is the probability of the domain classifier on the input, and a negative sign is added in front of it due to the adversarial training method.
In summary, the total loss function of YOLO-G in this paper is shown as follows:
(4)
where λ1, λ2, λ3 is set to 1.0, 0.5, and 0.05, respectively, α is the loss weights of the feature alignment branches, and through the ablation experiment, this paper sets α to 1.0.
Experiment
Datasets
Cityscapes [59] and Foggy-Cityscapes [60]: The two datasets contain the same number of image samples, of which 2975 training images and 500 validation images, the only difference is that the latter one is the dataset after adding fog through the rendering engine. To verify the excellent performance of the YOLO-G in cross-domain object detection, the maximum dense 0.02 is selected.
SIM 10K [61]: A dataset of vehicle driving scenes synthesized by the rendering engine of GTA5, containing 10,000 vehicle targets in various street scenes.
KITTI [62]: Autonomous driving dataset contains 7481 detailed annotated various targets.
BDD 100K [63]: Large-scale autonomous driving dataset, contains a variety of typical urban scenarios, this paper selects data under daytime conditions for experiments, including 36278 training images and 5258 validation images.
PASCAL VOC [64]: A large-scale object detection dataset, which contains 20 categories of detailed labeled images, we use VOC 2007 and VOC 2012 for experiment as in [29] with a total of 16551 images.
Clipart [45] and Watercolor: Just like VOC, both datasets contain 20 types of targets, containing 1000 and 2000 image samples, respectively.
As shown in Table 1, we summarize some characters of all the benchmark. It is clear that all the datasets have unique domain compared to each other, and this is why they are widely used in cross-domain object detection researches.
Experimental platform
The experiments are based on Ubuntu 18.04 LTS operating system, 16GB running memory, 1 Nividia RTX 3090 GPU with 24GB memory as the hardware platform, using PyTorch deep learning framework, Pycharm development platform, 11.7 Cuda, and Cudnn 8.0 acceleration environment.
Implementation details
Considering the size and category type of the training dataset, this paper selects YOLOV5-L as the baseline, in which the input image size is unified as 640×640, the training epoch is 200, and the warm-up period is 3 epochs, and the other parameters are consistent with SSDA-YOLO. The data augmentation strategy based on mosaic is adopted to improve the detection precision of the model for small targets. We quantified the effects of mosaic and SPPF in ablation experiments. The remaining unspecified parameter settings are consistent with the original YOLO model.
Evaluation metrics
The mean average precision (mAP) of the experiment is used as the evaluation index, and the IOU is specified to be 0.5. For the K→C, S→C experiment, since only the car target is detected, AP is used as the evaluation metric, and the threshold value of IOU is also set to 0.5.
Result
Detection from Cityscapes to Foggy Cityscapes.
As usual, we set Cityscapes as the source domain and Foggy-Cityscapes as the target domain. In such setting, we can test whether different cross-domain models can effectively eliminate the influence of fog when weather conditions change. It makes sense when the model extracts consistent features that are accurate enough to identify objects under fog occlusion conditions. In this paper, some of the state-of-the-art models are compared with YOLO-G, and the experimental results are shown in Table 2.
It can be seen from Table 2 that compared with the method based on the two-stage detection model, YOLO-G achieves outstanding detection precision of 47.8 mAP. Furthermore, even the baseline model gets 39.9 mAP, much higher than most semi-supervised cross-domain detection models. Benefits from the rich data enhancement, SPPF used by the YOLO model, we reach the same result with SOTA such as DSS, which takes a much more complex model structure and training strategy. Regarding the role of these tricks, we conduct detailed experiments and analyses in the ablation experiment section. Although the YOLO-G model and the DSS method both reach 47.8 mAP, the difference is that YOLO-G only adopts unsupervised self-learning mode. That is to say, we do not require a substantial improvement in model structure and training strategy, nor to train in stages. In short, YOLO-G gets the highest score in 5 categories, outstanding than others, and the detection results are shown in Fig 2.
Visual detection examples using the original YOLOv5 model: left column (source only), middle column (YOLO-G), right column (target only).
Detection from a different view.
Different image acquisition equipment and angles will cause very different target imaging results, which is a very common target cross-domain detection problem. To test the precision of YOLO-G object detection under such conditions, this paper uses SIM 10k, Cityscapes, and KITTI datasets with different perspectives for cross-experiments, because there is only a car target in SIM 10K, AP is used as an evaluation index for this section of the experiment, and the relevant results are presented in Table 3.
The advantages of YOLO-G in processing multi-view, multi-directional cross-domain scenes are obvious. Through the guidance of feature alignment branches, YOLO-G extracts the features of the target more accurately, ensuring that it can sample enough feature information about the target at any viewing angle, orientation, and scale to achieve effective detection. YOLO-G achieved a mAP of 64.2, and 62.8, which is at least 11.6 mAP better than the other two-stage models. However, in S→C, there is not only a difference in viewing angle but also a scene difference, YOLO-G only improves 0.2 mAP compared with the baseline model under the dual conditions of processing cross-angle and cross-virtual reality.
Detection from real to virtual.
In many cases, there is a big difference between the virtual scene and the real environment, which can easily lead to a large deviation from the model. Combined with the results of S→C experiments, this section tests the precision of YOLO-G in processing virtual and real scene conditions. The experiment sets the VOC dataset as the source domain, Clipart, and Watercolor as the target domains, and Watercolor has 6 types of targets consistent with the VOC dataset.
It can be seen from the comparison in Table 4 that compared with the two-stage detection models, YOLO-G achieves 44.3 mAP by using autonomous feature alignment in the cross-domain detection scenario of 20 categories. It is worth noting that SSDA-YOLO is also a cross-domain detection model based on YOLOV5-L, it adopts complex training strategies such as semi-supervision, EMA, knowledge distillation, consistency loss. Under semi-supervised conditions, SSDA-YOLO gets much higher mAP, but our YOLO-G is only equipped with GRL, and achieves a slightly higher mAP compared to TIA.
Table 5 shows that the YOLO-G model based on feature alignment constraints performs more prominently in the cross-virtual and real object detection tasks.
Detection from small dataset to large-scale dataset.
In this scenario, the experimental setting takes Cityscape as the source domain and the validation set of the large-scale autonomous driving dataset BDD 100k as the target domain. The performance of the proposed model in cross-domain detection between small sample dataset and large-scale dataset is fully tested. Because there are few train targets in BDD100k, therefore, the experimental process removes the category of the train, and the experimental results are in Table 6.
A small sample is an important challenge faced by deep learning models, avoiding overfitting problems is the core content, for the cross-domain transformation from small sample data to large-scale application scenarios, YOLO-G gives a good solution, Table 6 shows that under unsupervised learning conditions, the combination of GRL and YOLO can improve the generalization ability of the model. Its mAP reaches 34.6, but compared with the 53.4 mAP obtained by BDD100k large-scale training set, the YOLO-G model still has a lot of room for improvement.
Ablation experiments
Through the above experiments, the effectiveness of the YOLO-G model in cross-domain object detection is fully displayed. YOLO-G is based on YOLOV5, which has a series of tricks that are clearly different from two-stage detection models, such as Spatial Pyramid Pooling—Fast (SPPF), mosaic data augmentation. As we pointed out earlier, both tricks allow the baseline to gain excellent performance. At the same time, the distribution of pre-trained weights and weight coefficients will also affect the training results. In order to delve into the impact of each trick on YOLO-G, we conduct a full range of experiments to verify the authenticity of the model.
The impact of training tricks.
In this section, we mainly discuss 2 tricks used by the YOLO model, SPPF and mosaic. We make the model without SPPF and mosaic as vanilla model. Then, we verify the impact of these 2 tricks in the City → Foggy City scenario, and the relevant experi mental results are shown in Table 7.
It is clear from the Table 7 that the training effect of the Vanilla model is not ideal without any tricks, and it can be said that the performance is similar to that of DAF. With the addition of the trick, it can be seen that the detection accuracy of the model gradually increases, and SPPF increases the mAP by 0.5 mAP to 31 mAP. The mosaic increased the mAP by 0.8 mAP to 31.3 mAP. Finally, with both tricks turned on, we get that under source only, the YOLOV5-L model reaches 34.2 mAP, which is far more than the Faster RCNN model without these tricks.
Through the above experiment, it is suitable to use the YOLOV5-L model with various trick aids as the baseline, which is beneficial to improve the final detection accuracy of YOLO-G in source only setting.
The impact of pretrained-weight.
We mentioned early that with the pre-trained weights, the model has mastered some prior knowledge of the detection object. When comparing different models, the pre-trained weights loaded by different backbone models are not consistent. Therefore, it has a certain impact on the results. To this end, we carry out comparative experiments without the support of pre-trained weights in this part to test the self-learning ability of the YOLO-G model. To speed up the model training process, this set of experiments set VOC2012 as the source domain, and Clipart as the target domain, the model is carried out without loading the pre-training weights, and the rest of the parameter settings are consistent with the previous text, and the source only, YOLO-G, and target only comparison experiments are carried out respectively, and the relevant results are as shown in Table 8.
The pre-trained weights are the result of sufficient training in a very large-scale dataset, with the help of which the model can quickly extract the main features of the target in the initial stage, accelerating the speed of model convergence. Table 8 shows that the YOLO-G can quickly and accurately learn the main features of the target in cross-domain scenarios without the help of pre-trained weights. The loss curve of the training process is shown in Fig 3.
The impact of weight coefficients.
The training loss function of the YOLO-G model consists of two parts, to test the influence of feature alignment loss on the overall detection precision of the model, under the condition of using pre-training weights, the cross-domain object detection experiment from VOC2012 to Clipart is carried out, and the α is set to 0.1, 0.5, 1.0, 1.5, and 2.0 for group experiments, and the detection results of each category are shown in Table 9.
Comparative experiments show that the model achieves better detection precision when α is 1.0. The curve of mAP during the training process is shown in Fig 4. The left image shows that at a weight of 1, the detection accuracy can reach a higher level faster and more consistently. Therefore, the constraint weight of α of 1.0 is used for all the experiments.
Discuss and conclusion
YOLO-G gains much stronger ability in the cross-domain object detection, summarizing all these experiments. But there is a serious problem as well, YOLO-G shows a poor performance considering small objects. YOLO is an anchor-based model, so there exists conflict when deciding which anchor is much suitable for all the objects. Especially, when there are few small objects in the dataset, the model may be dominated by the larger and easily detected objects, without the help of focal loss. In summary, there is a lot to further explore heading for the real applications.
To alleviate the problem of cross-domain object detection, this paper analyzes the characteristics of mainstream algorithm models, and proposes a simple and efficient YOLO-G model based on the YOLOV5. By introducing feature alignment branch and adversarial training, we improve the consistency of the backbone model in extracting target features, enhance the generalization of the model, and achieve better cross-domain detection ability. We also organize 9 groups of cross-domain comparative experiments, and the YOLO-G model proposed in this paper achieves precision beyond a series of SOTA models, indicating that it has better application prospects in cross-domain object detection tasks.
References
- 1. Kou F, Du J, Cui W, Shi L, Cheng P, Chen J, et al. Common semantic representation method based on object attention and adversarial learning for cross-modal data in IoV. IEEE Transactions on Vehicular Technology 2019;68(12):11588–11598.
- 2. Shi L, Du J, Cheng G, Liu X, Xiong Z, Luo J. Cross‐media search method based on complementary attention and generative adversarial network for social networks. International Journal of Intelligent Systems 2022;37(8):4393–4416.
- 3. Shi L, Luo J, Zhu C, Kou F, Cheng G, Liu X. A survey on cross-media search based on user intention understanding in social networks. Information Fusion 2023;91:566–581.
- 4.
Redmon J, Farhadi A. YOLO9000: Better Faster Stronger. In: 2017 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2017; Honolulu, HI, USA: IEEE; 2017. p. 7263–7271.
- 5. Redmon J, Farhadi A. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767 2018.
- 6.
Redmon J, Divvala S, Girshick R, Farhadi A. You Only Look Once: Unified Real-Time Object Detection. In: 2016 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2016; Las Vegas, USA: IEEE; 2016. p. 779–788.
- 7. Bochkovskiy A, Wang C, Liao HM. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934 2020.
- 8.
Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C, et al. Ssd: Single shot multibox detector. In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14; 2016: Springer; 2016. p. 21–37.
- 9. He K, Gkioxari G, Dollár P, Girshick R. Mask r-cnn. In: Proceedings of the IEEE international conference on computer vision; 2017; 2017. p. 2961–2969.
- 10. Ren S, He K, Girshick R, Sun J. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 2015;28:91–99.
- 11.
Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S. End-to-end object detection with transformers. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16; 2020: Springer; 2020. p. 213–229.
- 12. Zhou H, Jiang F, Lu H. SSDA-YOLO: Semi-supervised domain adaptive YOLO for cross-domain object detection. Computer Vision and Image Understanding 2023:103649.
- 13.
Chen Y, Li W, Sakaridis C, Dai D, Van Gool L. Domain Adaptive Faster RCNN for Object Detection in the Wild. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2018; Salt Lake City, UT, USA: IEEE; 2018. p. 3339–3348.
- 14.
He Z, Zhang L. Domain Adaptive Object Detection via Asymmetric Tri-way Faster-RCNN. In: 2020 European Conference on Computer Vision(ECCV); 2020; Glasgow, UK: Springer; 2020.
- 15.
He M, Wang Y, Wu J, Wang Y, Li H, Bo L, et al. Cross Domain Object Detection by Target-Perceived Dual Branch Distillation. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022; New Orleans, LA, USA: IEEE; 2022. p. 9570–9579.
- 16.
Li J, Li G, Shi Y, Yu Y. Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021; Nashville, TN, USA: IEEE; 2021. p. 2505–2514.
- 17.
Lin C, Yuan Z, Zhao S, Sun P, Wang C, Cai J. Domain-Invariant Disentangled Network for Generalizable Object Detection. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021 2021; Nashville, TN, USA: IEEE; 2021. p. 8751–8760.
- 18.
Zheng Y, Huang D, Liu S, Wang Y. Cross-domain Object Detection through CoarsetoFine Feature Adaptation. In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020; Seattle, WA, USA: IEEE; 2020. p. 13766–13775.
- 19.
Regmi K, Shah M. Bridging the Domain Gap for Ground-to-Aerial Image Matching. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV); 2019; Seoul, Korea (South): IEEE; 2019. p. 470–479.
- 20. Shen Z, Maheshwari H, Yao W, Savvides M. Scl: Towards accurate domain adaptive object detection via gradient detach based stacked complementary losses. arXiv preprint arXiv:1911.02559 2019.
- 21. Nguyen D, Tseng W, Shuai H. Domain-adaptive object detection via uncertainty-aware distribution alignment. In: Proceedings of the 28th ACM international conference on multimedia; 2020; 2020. p. 2499–2507.
- 22.
Zhu X, Pang J, Yang C, Shi J, Lin D. Adapting Object Detectors via Selective Cross-Domain Alignment. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019; Long Beach, CA, USA: IEEE; 2019. p. 687–696.
- 23.
Rezaeianaran F, Shetty R, Aljundi R, Reino DO, Zhang S, Schiele B. Seeking Similarities Over Differences Similarity-Based Domain Alignment for Adaptive Object Detection. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV); 2021; Montreal, Canada: IEEE; 2021. p. 9204–9213.
- 24.
Saito K, Kim D, Sclaroff S, Darrell T, Saenko K. Semi-supervised Domain Adaptation via Minimax Entropy. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV); 2019; Seoul, Korea (South): IEEE; 2019. p. 8050–8058.
- 25.
Ramamonjison R, Dehkordi AB, Kang X, Bai X, Zhang Y. SimROD A Simple Adaptation Method for Robust Object Detection. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV); 2021; Nashville, TN, USA: IEEE; 2021. p. 3570–3579.
- 26. Csaba B, Qi X, Chaudhry A, Dokania P, Torr P. Multilevel knowledge transfer for cross-domain object detection. arXiv preprint arXiv:2108.00977 2021.
- 27. Ge Z, Liu S, Wang F, Li Z, Sun J. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430 2021.
- 28.
Hsu C, Tsai Y, Lin Y, Yang M. Every pixel matters: Center-aware feature alignment for domain adaptive object detector. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16; 2020: Springer; 2020. p. 733–748.
- 29.
Chen C, Zheng Z, Huang Y, Ding X, Yu Y. I3Net: Implicit Instance-Invariant Network for Adapting One-Stage Object Detectors. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021; Nashville, TN, USA: IEEE; 2021. p. 12576–12585.
- 30.
Li W, Liu X, Yuan Y. SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022; New Orleans, LA, USA: IEEE; 2022. p. 5291–5300.
- 31.
Zhou W, Du D, Zhang L, Luo T, Wu Y. Multi-Granularity Alignment Domain Adaptation for Object Detection. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022; New Orleans, LA, USA: IEEE; 2022. p. 9581–9590.
- 32. Vidit V, Salzmann M. Attention-based domain adaptation for single-stage detectors. Machine Vision and Applications 2022;33(5):65.
- 33. Liu W, Ren G, Yu R, Guo S, Zhu J, Zhang L. Image-adaptive YOLO for object detection in adverse weather conditions. In: Proceedings of the AAAI Conference on Artificial Intelligence; 2022; 2022. p. 1792–1800.
- 34.
Ganin Y, Lempitsky V. Unsupervised domain adaptation by backpropagation. In: International conference on machine learning; 2015: PMLR; 2015. p. 1180–1189.
- 35. Ganin Y, Ustinova E, Ajakan H, Germain P, Larochelle H, Laviolette F, et al. Domain-adversarial training of neural networks. The journal of machine learning research 2016;17(1):2096–2030.
- 36.
Li Y, Mao H, Girshick R, He K. Exploring plain vision transformer backbones for object detection. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX; 2022: Springer; 2022. p. 280–296.
- 37.
Li P, Li D, Li W, Gong S, Fu Y, Hospedales TM. A Simple Feature Augmentation for Domain Generalization. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV); 2021 2021; Nashville, TN, USA: IEEE; 2021. p. 8866–8875.
- 38.
Saito K, Ushiku Y, Harada T, Saenko K. Strong-Weak Distribution Alignment for Adaptive Object Detection. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019; Long Beach, CA, USA: IEEE; 2019. p. 6956–6965.
- 39.
Xu M, Wang H, Ni B, Tian Q, Zhang W. Cross-Domain Detection via Graph-Induced Prototype Alignment. In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020; Seattle, WA, USA: IEEE; 2020. p. 12355–12364.
- 40.
He Z, Zhang L. Multi-adversarial Faster-RCNN for Unrestricted Object Detection. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV); 2019; Seoul, Korea (South): IEEE; 2019. p. 6668–6677.
- 41. He Z, Zhang L, Yang Y, Gao X. Partial alignment for object detection in the wild. IEEE Transactions on Circuits and Systems for Video Technology 2021;32(8):5238–5251.
- 42.
Tian Z, Shen C, Chen H, He T. Fcos: Fully Convolutional One-Stage Object Detection. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV); 2019; Seoul, Korea (South): IEEE; 2019. p. 9627–9636.
- 43.
Park T, Efros AA, Zhang R, Zhu J. Contrastive learning for unpaired image-to-image translation. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16; 2020: Springer; 2020. p. 319–345.
- 44.
Zhang S, Tuo H, Hu J, Jing Z. Domain adaptive yolo for one-stage cross-domain detection. In: Asian Conference on Machine Learning; 2021: PMLR; 2021. p. 785–797.
- 45.
Inoue N, Furuta R, Yamasaki T, Aizawa K. Cross-Domain Weakly-Supervised Object Detection through Progressive Domain Adaptation. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2018; Salt Lake City, UT, USA: IEEE; 2018. p. 5001–5010.
- 46.
Cai Q, Pan Y, Ngo C, Tian X, Duan L, Yao T. Exploring Object Relation in Mean Teacher for CrossDomain Detection. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019; Long Beach, CA, USA: IEEE; 2019. p. 11457–11466.
- 47. Motiian S, Piccirilli M, Adjeroh DA, Doretto G. Unified deep supervised domain adaptation and generalization. In: Proceedings of the IEEE international conference on computer vision; 2017; 2017. p. 5715–5725.
- 48.
Deng J, Li W, Chen Y, Duan L. Unbiased Mean Teacher for Cross-domain Object Detection. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021; Nashville, TN, USA: IEEE; 2021. p. 4091–4101.
- 49.
Kim T, Jeong M, Kim S, Choi S, Kim C. Diversify and Match A Domain Adaptive Representation Learning Paradigm for Object Detection. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019; Long Beach, CA, USA: IEEE; 2019. p. 12456–12464.
- 50.
Xu CD, Zhao XR, Jin X, Wei XS. Exploring Categorical Regularization for Domain Adaptive Object Detection. In: 2020 European Conference on Computer Vision(ECCV); 2020; Glasgow, UK: Springer; 2020. p. 11724–11733.
- 51. Lin T, Dollár P, Girshick R, He K, Hariharan B, Belongie S. Feature pyramid networks for object detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition; 2017; 2017. p. 2117–2125.
- 52. Ben-David S, Blitzer J, Crammer K, Kulesza A, Pereira F, Vaughan JW. A theory of learning from different domains. Machine learning 2010;79:151–175.
- 53. Wang H, Liao S, Shao L. Afan: Augmented feature alignment network for cross-domain object detection. IEEE Transactions on Image Processing 2021;30:4046–4056. pmid:33793400
- 54.
Zhang X, Cao J, Shen C, You M. Self-Training With Progressive Augmentation for Unsupervised Cross-Domain Person Re-Identification. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV); 2019; Seoul, Korea (South): IEEE; 2019. p. 8222–8231.
- 55.
Chen C, Zheng Z, Ding X, Huang Y, Dou Q. Harmonizing Transferability and Discriminability for Adapting Object Detectors. In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020; Seattle, WA, USA: IEEE; 2020. p. 8869–8878.
- 56. Wang W, Cao Y, Zhang J, He F, Zha Z, Wen Y, et al. Exploring sequence feature alignment for domain adaptive detection transformers. In: Proceedings of the 29th ACM International Conference on Multimedia; 2021; 2021. p. 1730–1738.
- 57.
Wang Y, Zhang R, Zhang S, Li M, Xia Y, Zhang X, et al. Domain-Specific Suppression for Adaptive Object Detection. In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021; Montreal, Canada: IEEE; 2021. p. 9603–9612.
- 58.
Zhao L, Wang L. Task-specific Inconsistency Alignment for Domain Adaptive Object Detection. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022; New Orleans, LA, USA: IEEE; 2022. p. 14217–14226.
- 59.
Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, et al. The Cityscapes Dataset for Semantic Urban Scene Understanding. In: 2017 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2016; Las Vegas, USA: IEEE; 2016. p. 3213–3222.
- 60. Sakaridis C, Dai D, Van Gool L. Semantic foggy scene understanding with synthetic data. International Journal of Computer Vision 2018;126:973–992.
- 61. Johnson-Roberson M, Barto C, Mehta R, Sridhar SN, Rosaen K, Vasudevan R. Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks? arXiv preprint arXiv:1610.01983 2016.
- 62.
Geiger A, Lenz P, Urtasun R. Are we ready for autonomous driving? the kitti vision benchmark suite. In: 2012 IEEE conference on computer vision and pattern recognition; 2012: IEEE; 2012. p. 3354–3361.
- 63. Yu F, Chen H, Wang X, Xian W, Chen Y, Liu F, et al. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition; 2020; 2020. p. 2636–2645.
- 64. Everingham M, Van Gool L, Williams CK, Winn J, Zisserman A. The pascal visual object classes (voc) challenge. International journal of computer vision 2009;88:303–308.