Fig 1.
Comparison of parameters and cost of different detection models.
(a) Model Parameters. (b) Computational Cost (FLOPs).
Fig 2.
The structure of the proposed DOLD-Net model.
The DBOAN backbone extracts hierarchical multi-scale features (B2–B5), the CGFP integrates and fuses them through the CSFE, FSAM, and CAOFM to generate the feature pyramid (P2–P5), and the Decoder performs object classification and bounding box prediction.
Fig 3.
The structure of DBOB.
Fig 4.
The structure of CSFE.
Fig 5.
Image sourced from [20] under the French Open License 2.0 license.
Fig 6.
The structure of FSAM.
Fig 7.
The structure of CAOFM.
Table 1.
Performance comparison of different models on the dataset.
Table 2.
Performance comparison of different models on the dataset.
Table 3.
Performance comparison of different models on the CherryChèvre dataset.
Table 4.
Performance comparison of different models on the SheepCounter dataset.
Table 5.
Performance comparison of different models on the ChickenFlow dataset.
Table 6.
Ablation results of different modules across various datasets.
Table 7.
Ablation results of different CGFP module components.
Table 8.
Ablation results of different kernel size settings.
Table 9.
Impact of different propagation counts and locations on model performance.
Table 10.
Impact of different fusion methods on the performance of the model.
Table 11.
Impact of different feature extraction modules on the performance of the model.
Table 12.
Comparison of different wavelet decomposition methods in the FSAM module.
Table 13.
Impact of different core operators within the CSFE module on model performance.
Fig 8.
Visual results of different models on the task of livestock detection under dense occlusion.
Green boxes denote correctly detected targets (TP), red boxes denote falsely detected targets (FP), and blue boxes denote missed targets (FN). Images in rows 1, 4, 5 and 8 sourced from [19], available under the CC0 public domain dedication. Images in rows 2 and 6 sourced from [43] under the CC BY 4.0 license. Images in rows 3 and 7 sourced from [3] under the CC BY 4.0 license.
Fig 9.
Heatmap across different datasets.
Warmer colors indicate stronger model responses or attention, whereas cooler colors indicate weaker responses. Images in rows 1, 2 and 3 sourced from [19], available under the CC0 public domain dedication. Images in rows 4 sourced from [3] under a CC BY 4.0 license. Images in rows 5 and 6 sourced from [43] under the CC BY 4.0 license.
Fig 10.
Feature map heatmaps with and without CSFE.
Image sourced from [19], available under the CC0 public domain dedication.
Fig 11.
Failure case of DOLD-Net under complex pose conditions.
Image sourced from [19], available under the CC0 public domain dedication.
Fig 12.
Comparison of 1D Semantic Sequence Unrolling between Normal and Extreme Poses.
Image sourced from [19], available under the CC0 public domain dedication.