Papers
arxiv:2610.01762

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

Published on Oct 1
· Submitted by
Xiangyu Zeng
on Oct 2
#1 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.

Community

Paper author Paper submitter

Hi everyone! We’re sharing OneStreamer, a 4B model that unifies perception, memory, and proactive responses for streaming video interaction.

The core idea is to learn both what to remember and when to respond. OneStreamer records evidence as video unfolds, before a future question is known, and responds when sufficient evidence becomes available.

Key highlights:

  • Memory beyond the recent visual context: Proactive Hierarchical Caption Memory (PHCM) turns observations into timestamped captions and event summaries, preserving evidence for later questions after the original frames leave the visual window.
  • Learning when to respond: Proactive State Transition Learning (PSTL) outperforms dense state supervision while supervising only 27.5% of annotated state tokens.
  • Strong results at 4B: OneStreamer achieves the best aggregate results among the compared methods on all eight evaluated benchmarks, covering perception, memory, and proactive response.
  • OneStreamer-1M: We introduce a dataset with over one million streaming video interaction records, combining synthesized captions and QA with cleaned existing data.

📺 Project page and video demo
💻 Code

We’re happy to discuss the method, data, and evaluation, and would love to hear your feedback!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.01762
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 1

Spaces citing this paper 1

Collections including this paper 2