Streaming models
Continuous video, audio, memory, and real-time interaction.
Streaming Proactivity
README · Applications · Personal Agents · Projects & Products · Research Map · Benchmark Matrix
Streaming proactivity is a dedicated cross-cutting module: it connects models, inference frameworks, memory, training data, and evaluation. Paper entries remain in the bibliography. This guide asks who chooses the next output and what evidence can cause that choice.
Last checked: 2026-10-04. Mechanisms and availability are based on primary papers and official repositories. Models were not run as part of this curation, and scores from different protocols are not ranked together.
Capability Routes
These routes can overlap. They describe a behavior, not an increasing capability ranking.
| Route | Decision being made | Representative resources | Evidence boundary |
|---|---|---|---|
| Scene-triggered intervention | Does an unfolding event merit speech, an alert, guidance, or continued silence? | JoyAI-VL-Interaction, MOSS-VL-Realtime, Proact-VL | Check the standing instruction, scene trigger, and negative/silent cases. |
| Evidence-conditioned response timing | Has enough evidence arrived to answer a standing question? | StreamReady, Eyes Wide Open, StreamOV, Response-G1; OneStreamer's QA branch | This establishes wait/respond behavior; discovering a new user need requires additional evidence. |
| Proactive evidence recording | What should be recorded before a future query is known? | OneStreamer's caption-memory branch | Internal memory generation and user-visible intervention are different outputs. |
| Full-duplex interaction and delegation | When should the agent listen, speak, interrupt, or return from background work while perception continues? | MiniCPM-o 4.5, Gander; JoyAI's delegation branch | Full-duplex execution needs content-grounded initiative, not only the ability to overlap input and output. |
Models and Frameworks
Scene-Triggered Interaction
| Work / source | Mechanism and proactive signal | Resources / release snapshot | Scope boundary |
|---|---|---|---|
| JoyAI-VL-Interaction · arXiv 2606 | 8B vision-first model chooses silent / respond / delegate each second. AdaCodec compresses predictable frames; ASR/TTS, memory, and background reasoning are pluggable. | Paper · Code · Model · Data · Notes. Public model/data/system links. | Demonstrated scenarios and human preferences do not establish general computer-use autonomy or broad superiority. |
| MOSS-VL-Realtime · arXiv 2608 | Continuous video perception and generation with model-controlled speaking/silence. Cross-attention supports incrementally arriving visual context. | Paper · Code · Model · Notes. Realtime inference, serving/demo, and SFT resources are available; full training engine remains a roadmap item. | Select the Realtime variant; offline Instruct and Base variants have different intended interfaces. |
| Proact-VL · ICML 2026 per arXiv record | Streaming companion framework learns when to respond in solo commentary, co-commentary, and user guidance. | Paper · Code / model zoo / data · Notes. Concrete checkpoint/data/evaluation links appear below an outdated TODO block. | Gaming commentary/guidance is a useful intervention setting, with distinct task instructions and speech expectations. |
Evidence, Memory, and Response Timing
| Work / source | Mechanism and proactive signal | Resources / release snapshot | Scope boundary |
|---|---|---|---|
| OneStreamer · arXiv 2610.01762 | PHCM records query-independent detail/event captions; PSTL learns state transitions. QA states include Silence, Standby, and Response. |
Paper · Code · Model card · Data · Notes. Hugging Face publicly lists ungated weight files; the repository availability table still says pending (conflicting snapshot). | Distinguish internal caption generation from externally useful initiative. The paper evaluates perception, memory, and proactive-response benchmarks. |
| StreamReady · CVPR 2026 | Answer-readiness supervision learns when enough streaming evidence has arrived; premature and delayed answers receive asymmetric penalties. | Paper · Project · Notes. Resource entry is the project page. | A query-conditioned readiness policy, evaluated with ProReady-QA; do not equate it with independent need discovery. |
| StreamOV · arXiv 2605 | Evidence-guided long/short-term multimodal memory under a fixed budget; hidden-state trigger chooses when to respond without generating explicit silence tokens. | Paper · Notes. This curation verified the paper, not a canonical public implementation/checkpoint. | SOVBench-T tests queried evidence appearing versus never appearing; SOVBench-O also tests multi-turn contextual inference. |
| Response-G1 · arXiv 2605; repository reports ACL 2026 | Fine-tuning-free, query-guided scene graphs → historical graph retrieval → silence / response trigger prompting. | Paper · Code · Notes. Official evaluation scripts are available; uses existing Video-LLM backbones. | An inference framework with explicit evidence, rather than a newly trained interaction foundation model. |
| Eyes Wide Open (EyeWO) · NeurIPS 2025 per arXiv record | Egocentric streaming answers at opportune moments, using a data engine, multi-stage training, and proactive dynamic compression. | Paper · Code · Notes. Training/evaluation code and ModelScope resource identifiers are documented. | ESTP-Bench emphasizes coherent, timely answers to evolving questions; repository setup contains placeholders that need inspection. |
Full-Duplex Interaction and Agent Delegation
| Work / source | Mechanism and proactive signal | Resources / release snapshot | Scope boundary |
|---|---|---|---|
| MiniCPM-o 4.5 · arXiv 2604 | 9B omni-modal model; Omni-Flow aligns simultaneous audio/video input and speech/text output. Supports scene-conditioned reminders/comments. | Paper · Code · Model · Notes. Model and live-interaction resources are public. | Select MiniCPM-o 4.5, not a newer vision-only MiniCPM-V release. Overlapping I/O and intervention quality need separate evaluation. |
| Gander / Omni Interaction Agent · arXiv 2609 | Learned listen / speak control connects a streaming Thinker/Talker to an asynchronous reasoning/tool-use Brain. | Paper · Code · Models · Notes. Runtime/models linked; README says dataset release is pending. | Separate interaction model, speech generation, background agent, and long-session context policy when comparing systems. |
Evaluation Routes
Select a benchmark by the decision it measures, not simply whether its inputs are video. Mixed benchmarks require reporting the relevant proactive subset.
| Evaluation target | Representative benchmark | What it can establish | Primary entry |
|---|---|---|---|
| Egocentric response coherence and timing | ESTP-Bench / Eyes Wide Open | Timely responses to evolving questions; ESTP-F1 combines the task's proactive criteria. | Paper · Code |
| Gaze-conditioned intention and alerting | StreamGaze | Past/present temporal reasoning and two proactive tasks, using egocentric video plus gaze. A benchmark, not a new model. | Paper · Code · Notes |
| Answer readiness | ProReady-QA / StreamReady | Correctness relative to annotated evidence windows and early/late answer cost. | Project |
| Multi-turn omni-video and absent-evidence silence | SOVBench-O / SOVBench-T | Multi-turn QA plus positive/negative response-trigger decisions, including precision, recall, and F1. | Paper / protocol |
| Live commentary and guidance | Live Gaming Benchmark / Proact-VL | Solo/co-commentary and user guidance; content quality, timing, and response frequency metrics. | Code / protocol · Data |
| Multimodal interaction and scene alerts | OmniMMI, OmniPro | Proactive alert/output tasks alongside other streaming interaction tasks. | OmniMMI · OmniPro |
| Personalized egocentric assistance | EgoPro-Bench, EgoServe | User-conditioned triggers, silence, and continuous assistance. | EgoPro-Bench · EgoServe |
| Duplex initiative and content-grounded floor control | DuplexAct-Bench, Full-Duplex Floor Selection | Proactive initiation and active silence; whether an opportunity to speak is also a reason to intervene. | See source records in the Benchmark Matrix. |
| Streaming execution validity | TRACE | Audits how an evaluation actually supplies input and runs inference. | See its source/protocol entry in the Benchmark Matrix. |
For new experiments, report causal input exposure, standing instructions, intervention positives and silence negatives, timing error, false alerts, content utility, and the latency boundary. If historical frames or future evidence are visible, disclose that protocol rather than calling the result a live intervention test. Separate caption-memory construction cost from answer-generation cost.
Adjacent Infrastructure and Discovery Lists
- Live VLM WebUI: camera/serving interface; proactive policy depends on the connected backend.
- Awesome VLM Streaming Video: discovery by streaming interaction, memory, models, and evaluation.
- Awesome Streaming Video Understanding: model/paper/dataset/benchmark discovery.
- Awesome Streaming Agents: continuous-input and active-activation perspective.
Discovery lists help find candidates. Individual mechanism and release claims in this guide are checked against primary sources; unresolved resources are identified in the matrix above.