OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
The authors introduce OneStreamer, an approach that jointly learns query-independent evidence recording and task response through proactive generation for streaming video. It utilizes Proactive Hierarchical Caption Memory (PHCM) to produce local-detail captions and event summaries. Combined with the OneStreamer-1M dataset containing over one million records, their 4B model outperforms other methods across eight streaming video understanding benchmarks while supervising only 27.5% of annotated state tokens.
OneStreamer-1M is a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks.
Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks.
PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens.