CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 202 upvotes

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

QUESTION — How can reusable factual memory be maintained in streaming video interaction without compromising real-time perception?

The authors introduce OneStreamer, an approach that jointly learns query-independent evidence recording and task response through proactive generation for streaming video. It utilizes Proactive Hierarchical Caption Memory (PHCM) to produce local-detail captions and event summaries. Combined with the OneStreamer-1M dataset containing over one million records, their 4B model outperforms other methods across eight streaming video understanding benchmarks while supervising only 27.5% of annotated state tokens.

OneStreamer-1M is a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks.

Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks.

PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens.

Lanxingxuan · 01 Oct 2026 read the original ↗
↑