Streaming models

Continuous video, audio, memory, and real-time interaction.

37 papers

Streaming Proactivity

README · Applications · Personal Agents · Projects & Products · Research Map · Benchmark Matrix

Streaming proactivity is a dedicated cross-cutting module: it connects models, inference frameworks, memory, training data, and evaluation. Paper entries remain in the bibliography. This guide asks who chooses the next output and what evidence can cause that choice.

Last checked: 2026-10-04. Mechanisms and availability are based on primary papers and official repositories. Models were not run as part of this curation, and scores from different protocols are not ranked together.

Capability Routes

These routes can overlap. They describe a behavior, not an increasing capability ranking.

Route Decision being made Representative resources Evidence boundary
Scene-triggered intervention Does an unfolding event merit speech, an alert, guidance, or continued silence? JoyAI-VL-Interaction, MOSS-VL-Realtime, Proact-VL Check the standing instruction, scene trigger, and negative/silent cases.
Evidence-conditioned response timing Has enough evidence arrived to answer a standing question? StreamReady, Eyes Wide Open, StreamOV, Response-G1; OneStreamer's QA branch This establishes wait/respond behavior; discovering a new user need requires additional evidence.
Proactive evidence recording What should be recorded before a future query is known? OneStreamer's caption-memory branch Internal memory generation and user-visible intervention are different outputs.
Full-duplex interaction and delegation When should the agent listen, speak, interrupt, or return from background work while perception continues? MiniCPM-o 4.5, Gander; JoyAI's delegation branch Full-duplex execution needs content-grounded initiative, not only the ability to overlap input and output.

Models and Frameworks

Scene-Triggered Interaction

Work / source Mechanism and proactive signal Resources / release snapshot Scope boundary
JoyAI-VL-Interaction · arXiv 2606 8B vision-first model chooses silent / respond / delegate each second. AdaCodec compresses predictable frames; ASR/TTS, memory, and background reasoning are pluggable. Paper · Code · Model · Data · Notes. Public model/data/system links. Demonstrated scenarios and human preferences do not establish general computer-use autonomy or broad superiority.
MOSS-VL-Realtime · arXiv 2608 Continuous video perception and generation with model-controlled speaking/silence. Cross-attention supports incrementally arriving visual context. Paper · Code · Model · Notes. Realtime inference, serving/demo, and SFT resources are available; full training engine remains a roadmap item. Select the Realtime variant; offline Instruct and Base variants have different intended interfaces.
Proact-VL · ICML 2026 per arXiv record Streaming companion framework learns when to respond in solo commentary, co-commentary, and user guidance. Paper · Code / model zoo / data · Notes. Concrete checkpoint/data/evaluation links appear below an outdated TODO block. Gaming commentary/guidance is a useful intervention setting, with distinct task instructions and speech expectations.

Evidence, Memory, and Response Timing

Work / source Mechanism and proactive signal Resources / release snapshot Scope boundary
OneStreamer · arXiv 2610.01762 PHCM records query-independent detail/event captions; PSTL learns state transitions. QA states include Silence, Standby, and Response. Paper · Code · Model card · Data · Notes. Hugging Face publicly lists ungated weight files; the repository availability table still says pending (conflicting snapshot). Distinguish internal caption generation from externally useful initiative. The paper evaluates perception, memory, and proactive-response benchmarks.
StreamReady · CVPR 2026 Answer-readiness supervision learns when enough streaming evidence has arrived; premature and delayed answers receive asymmetric penalties. Paper · Project · Notes. Resource entry is the project page. A query-conditioned readiness policy, evaluated with ProReady-QA; do not equate it with independent need discovery.
StreamOV · arXiv 2605 Evidence-guided long/short-term multimodal memory under a fixed budget; hidden-state trigger chooses when to respond without generating explicit silence tokens. Paper · Notes. This curation verified the paper, not a canonical public implementation/checkpoint. SOVBench-T tests queried evidence appearing versus never appearing; SOVBench-O also tests multi-turn contextual inference.
Response-G1 · arXiv 2605; repository reports ACL 2026 Fine-tuning-free, query-guided scene graphs → historical graph retrieval → silence / response trigger prompting. Paper · Code · Notes. Official evaluation scripts are available; uses existing Video-LLM backbones. An inference framework with explicit evidence, rather than a newly trained interaction foundation model.
Eyes Wide Open (EyeWO) · NeurIPS 2025 per arXiv record Egocentric streaming answers at opportune moments, using a data engine, multi-stage training, and proactive dynamic compression. Paper · Code · Notes. Training/evaluation code and ModelScope resource identifiers are documented. ESTP-Bench emphasizes coherent, timely answers to evolving questions; repository setup contains placeholders that need inspection.

Full-Duplex Interaction and Agent Delegation

Work / source Mechanism and proactive signal Resources / release snapshot Scope boundary
MiniCPM-o 4.5 · arXiv 2604 9B omni-modal model; Omni-Flow aligns simultaneous audio/video input and speech/text output. Supports scene-conditioned reminders/comments. Paper · Code · Model · Notes. Model and live-interaction resources are public. Select MiniCPM-o 4.5, not a newer vision-only MiniCPM-V release. Overlapping I/O and intervention quality need separate evaluation.
Gander / Omni Interaction Agent · arXiv 2609 Learned listen / speak control connects a streaming Thinker/Talker to an asynchronous reasoning/tool-use Brain. Paper · Code · Models · Notes. Runtime/models linked; README says dataset release is pending. Separate interaction model, speech generation, background agent, and long-session context policy when comparing systems.

Evaluation Routes

Select a benchmark by the decision it measures, not simply whether its inputs are video. Mixed benchmarks require reporting the relevant proactive subset.

Evaluation target Representative benchmark What it can establish Primary entry
Egocentric response coherence and timing ESTP-Bench / Eyes Wide Open Timely responses to evolving questions; ESTP-F1 combines the task's proactive criteria. Paper · Code
Gaze-conditioned intention and alerting StreamGaze Past/present temporal reasoning and two proactive tasks, using egocentric video plus gaze. A benchmark, not a new model. Paper · Code · Notes
Answer readiness ProReady-QA / StreamReady Correctness relative to annotated evidence windows and early/late answer cost. Project
Multi-turn omni-video and absent-evidence silence SOVBench-O / SOVBench-T Multi-turn QA plus positive/negative response-trigger decisions, including precision, recall, and F1. Paper / protocol
Live commentary and guidance Live Gaming Benchmark / Proact-VL Solo/co-commentary and user guidance; content quality, timing, and response frequency metrics. Code / protocol · Data
Multimodal interaction and scene alerts OmniMMI, OmniPro Proactive alert/output tasks alongside other streaming interaction tasks. OmniMMI · OmniPro
Personalized egocentric assistance EgoPro-Bench, EgoServe User-conditioned triggers, silence, and continuous assistance. EgoPro-Bench · EgoServe
Duplex initiative and content-grounded floor control DuplexAct-Bench, Full-Duplex Floor Selection Proactive initiation and active silence; whether an opportunity to speak is also a reason to intervene. See source records in the Benchmark Matrix.
Streaming execution validity TRACE Audits how an evaluation actually supplies input and runs inference. See its source/protocol entry in the Benchmark Matrix.

For new experiments, report causal input exposure, standing instructions, intervention positives and silence negatives, timing error, false alerts, content utility, and the latency boundary. If historical frames or future evidence are visible, disclose that protocol rather than calling the result a live intervention test. Separate caption-memory construction cost from answer-generation cost.

Adjacent Infrastructure and Discovery Lists

Discovery lists help find candidates. Individual mechanism and release claims in this guide are checked against primary sources; unresolved resources are identified in the matrix above.