Interaction Tracking Model
Establishing interaction tracking as a first-class perception target
We built Palona’s first lightweight interaction-tracking model, establishing interaction tracking as a first-class perception target. Compared with a flagship VLM baseline, the model raised accuracy from under 50% to over 80% while reducing per-frame latency from 15 seconds to 0.22 seconds.
The next challenge is turning this capability into a reliable customer solution. Customers may provide abundant video, but little or no labeled data for the interactions that matter in their environments. Our work explores how grounding perception in objects and actors, together with spatial relationships, temporal continuity, semantic context, and self-supervised or JEPA-style learning, can reduce annotation requirements and accelerate the path from raw customer video to deployable interaction models. This is an ongoing research and engineering effort, with more details to come.