SCIEPublish

A Small-Object Detection Model Based on Multi-Module Collaborative Optimization

Article Open Access

A Small-Object Detection Model Based on Multi-Module Collaborative Optimization

Author Information
College of Applied Mathematics, Chengdu University of Information Technology, Chengdu 610225, China
*
Authors to whom correspondence should be addressed.

Received: 06 July 2026 Revised: 24 August 2026 Accepted: 11 September 2026 Published: 30 September 2026

Creative Commons

© 2026 The authors. This is an open access article under the Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/).

Views:65
Downloads:30
Drones Auton. Veh. 2026, 3(4), 10024; DOI: 10.70322/dav.2026.10024
ABSTRACT: Small object detection in complex scenes, particularly for Unmanned Aerial Vehicle (UAV) aerial imagery, remains a core challenge in computer vision, primarily due to scarce feature information, vulnerability to background interference, high sensitivity of the traditional intersection over union (IoU) metric to minor positional deviations, and severe background clutter. To address these issues, this paper proposes a small-object detection model based on multi-module collaborative optimization. Built upon the YOLOv8 framework, the model introduces systematic improvements in three aspects: feature fusion, feature enhancement, and loss function design. Specifically, a weighted bidirectional feature pyramid network (BiFPN) is incorporated to provide richer multi-scale contextual information for small objects; a novel global-local collaborative attention mechanism (GLSA) is embedded to enhance the discriminative power of key features; and a bounding box regression loss based on Wasserstein distance is adopted to deliver smoother gradient behavior during training. The proposed model is comprehensively evaluated on a public dataset comprising numerous aerial and ground small objects. Compared with the baseline YOLOv8, our model achieves substantial performance gains: the optimal F1-confidence threshold rises from 0.42 to 0.67, while the mean average precision at IoU = 0.5 (mAP@0.5) increases from 81.4% to 81.8%, yielding a 0.4 percentage point improvement. These results indicate that the predicted confidences are better calibrated with respect to true accuracy, thereby enhancing the reliability of output detections. Meanwhile, the F1-score, which balances precision and recall, improves from 77% to 78%. Ablation studies further confirm that BiFPN and GLSA jointly improve classification and recall for small objects, whereas the Wasserstein Loss specifically optimizes localization accuracy. Together, these three modules constitute a cohesive high-performance detection pipeline. This work provides an effective and reliable solution for real-time small object detection in complex environments.
Keywords: YOLOv8; BiFPN; GLSA; Wasserstein Loss; Small object detection
TOP