Structured Evidence and Vision-Language Models for Interpretable Vision-Only UAV Behavior Analysis

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Vision-only counter-unmanned aerial vehicle (counter-UAV) systems offer low deployment cost, passive sensing, and compatibility with existing camera infrastructure, but most existing pipelines are limited to detection and tracking. As a result, operators receive target locations but little auditable evidence about the UAV's behavior or the urgency of the required response. This paper presents a four-layer framework for interpretable vision-only UAV behavior analysis. The framework first detects UAVs with a single-class YOLOv8 detector, associates detections with ByteTrack, repairs fragmented trajectories through greedy track merging, and converts each trajectory into structured motion evidence including speed statistics, linearity, scale-change ratio, curvature, and hovering ratio. The evidence is then combined with a compact key-frame mosaic and provided to a vision-language model, which produces behavior labels, response-urgency estimates, and natural-language rationales. Experiments on 20 DUT Anti-UAV sequences produced 52 valid tracks, of which 50 were manually annotated for behavior evaluation. The full multimodal configuration achieved the best primary-match accuracy (0.700) and any-match accuracy (0.760), while an image-only variant achieved only 0.220 primary-match accuracy. The results show that structured trajectory evidence is the dominant information source for UAV behavior recognition, while image evidence provides a smaller but useful gain in geometrically ambiguous cases.

Article activity feed