BehaviorScope-X: reusing pose-trained visual representations for full-video ethology

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Pose-estimation pipelines usually export keypoint coordinates and discard the intermediate visual representations learned to localize animals in a specific assay. We asked whether those discarded representations can be reused for full-video ethology. BehaviorScope-X tests this idea by treating a trained pose checkpoint as both a keypoint estimator and a reusable visual encoder: the pose model is run once to cache detections, keypoints, pose-derived social geometry, and frozen intermediate descriptors, after which compact temporal classifiers are trained on cached multimodal windows. Across MARS resident-intruder videos, cached pose-trained descriptors and pose-derived geometry provided complementary evidence for behavior decoding, recovering sustained behavioral episodes and local sequence structure while revealing a main limitation in dense short-bout regions. The same cache-and-classify design generalized across pose routes, including a MobileNetV3 backbone and a DeepLabCut SuperAnimal HRNet-W32 checkpoint, showing that standard pose workflows can expose behavior-relevant visual descriptors without giving up their keypoint-estimation role. We further tested the approach in Fly-v-Fly aggression, extending the analysis to a second species and shorter behavioral time scale, where sub-second events and annotation-boundary uncertainty limited strict bout recovery. End-to-end profiling showed that the workflow can operate near-real-time or real-time on consumer hardware. Together, these experiments support amortized pose vision as a practical strategy for reusing assay-trained pose models as stable sources of visual and geometric evidence for scalable behavioral analysis.

Author summary

Pose-estimation models are usually used to convert animal video into body landmarks, while the same model’s internal visual representations are discarded. We ask whether the model that estimates pose can also provide visual features for behavior analysis. Our workflow runs a trained pose model once, caches its landmarks, pose-derived interaction geometry, and internal visual features, and trains compact behavior classifiers on that cache. Across multiple pose backbones and pose-estimation workflows, these shared pose-and-visual signals supported full-video ethogram recovery without training a separate video network. This turns pose training into a reusable source of assay-specific visual and geometric evidence for scalable behavioral analysis.

Highlights

  • A pose-estimation checkpoint is reused as both keypoint estimator and visual encoder.

  • A single pose-inference pass caches keypoints, social geometry, and intermediate visual descriptors.

  • Cached pose-derived visual and geometric evidence improves full-video behavior decoding beyond pose-derived features alone.

  • BehaviorScope-X turns existing pose workflows into reusable full-video behavior-analysis pipelines.

Article activity feed