Benchmarks
Datasets and evaluation suites for proactive agents.
/
ContextClarify
Clarify or Answer
RealHumanEval
Need Help?
ProactiveBench
Proactive Agent
ESTP-Bench
Eyes Wide Open
SOVBench-O / SOVBench-T
StreamOV
Live Gaming Benchmark
Proact-VL
Full-Duplex Floor Selection
Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
ProAction
Cognitive Action Reasoning for Proactive Robots
ContextAgentBench
ContextAgent
PROBE
Beyond Reactivity
UserVille
Training Proactive and Personalized LLM Agents
ClarifyBench
Structured Uncertainty guided Clarification
ChronosBench
Long-term Task-oriented Agent
ProactiveBench (MLLM)
ProactiveBench / Trento
Pare-Bench
Pare
LatentNeeds-Bench
PASK
CogEval-Bench
CogniFold
ProCodeBench
Proactive Coding Assistants
ProactBench
Beyond What The User Asked For
ProActEval
Anticipate and Learn
GuidanceSalesBench
See, Infer, Intervene
Int-Bench
AI Assistants Overassist
AskBench
When and What to Ask
Act2Intention Bench
Act2Intention
Interactive Visual Grounding
When Seeing Is Not Enough
ProReady-QA
StreamReady
PROS-Bench
Beyond Instruction-Driven Editing
CONFLICTGUI
Do GUI Agents Know When Not to Act?
OR-Clarify
Ask Before You Optimize
TIMELI
Time-Aware Assistive Navigation
Physical Experiment Selection
New Evidence, Same Choice
IntentFlux
When Users Change Their Minds
PROACTIVITY-GYM
Foundations of Proactive Agents
Drift-Bench++
Beyond Oracle Communication
SWE-Intervene
Learning When and How to Intervene
Benchmark Matrix
This page compares benchmarks by what they actually test. The goal is to make benchmark selection faster than scanning individual paper summaries.
Layered Must Read · Streaming models and evaluation routes · Projects & Products
Quick Takeaways
- Best long-horizon personal assistant benchmarks: VibeLifeBench, π-Bench, VitaBench 2.0, ChronosBench.
- Best computer-use benchmarks: Act2Intention Bench, Ambig-SWE, CONFLICTGUI, ProAgentBench, PIRA-Bench, GUIDE, KnowU-Bench, Pare-Bench, ProCodeBench.
- Best dialogue, clarification, and premise-critique benchmarks: AskBench, CC-Mediation, ClarifyBench, ContextClarify, Drift-Bench++, FinInteract, IdeaAMBIG, IntentFlux, Interactive Visual Grounding, OR-Clarify, ProactBench, ProactiveEval, ProMediConv, ProMISe, RPCBench, Ψ-Bench.
- Best memory-oriented benchmarks: APM-Bench, CogEval-Bench, MemEye, TWIST, VitaBench 2.0.
- Best human-factor / timing benchmarks: DuplexAct-Bench, Int-Bench, JarvisBench, PASTABench, PROACTIVITY-GYM, RealHumanEval, Pare-Bench, ProAgentBench, ProactiveBench (MLLM).
- Streaming multimodal benchmarks by task (see evaluation routes): ESTP-Bench, StreamGaze, SOVBench, Live Gaming Benchmark, APM-Bench, DuplexAct-Bench, Full-Duplex Floor Selection, Live Assistant, TRACE, TIMELI, ProReady-QA, OmniAssistBench, OmniPro, EgoPro-Bench, OmniMMI, StreamArena, IPIBench, EgoServe, StreamSoccer.
Benchmark Matrix
| Benchmark | Paper | Domain | Input Stream | Proactive Target | User Model | Data Type | Main Metrics / Signals | Code / Data |
|---|---|---|---|---|---|---|---|---|
| ProMISe | ProMISe | Information-seeking dialogue | Multi-turn dialogue context | Ask proactive clarification questions | Simulated or annotated intent ambiguity | Dataset | Intent resolution, clarification quality | ACL |
| ContextClarify | Clarify or Answer | Visual question answering | Image-question pair plus optional external context reply | Decide whether to ask or answer, then ask one focused question | Human-verified missing context and non-ambiguous contrast cases | Ambiguous + contrast-set benchmark | Ask-versus-answer decision, clarification quality, end-to-end VQA accuracy | arXiv |
| FinInteract | FinInteract | Financial question answering | Ambiguous English or Chinese question, regulatory filing, and clarification reply | Elicit the intended interpretation and integrate it instead of choosing the default reading | Intended and default interpretations for each question | 173 bilingual ambiguity instances | Clarification elicitation, targeting, intent integration, default-versus-intended accuracy | arXiv |
| RealHumanEval | Need Help? | Programming | IDE task state and code context | Offer proactive programming help | Human participants | Human study | Completion, acceptance, disruption, preference | arXiv |
| Ambig-SWE | Ambig-SWE | Software engineering | Repository, underspecified issue, and interactive user responses | Detect underspecificity, ask targeted questions, and use the answer | Simulator backed by withheld issue information | SWE-bench Verified-derived interactive benchmark | Detection, clarification quality, post-interaction task resolution | OpenReview · GitHub |
| ProactiveBench | Proactive Agent | Desktop activity | Event streams and desktop context | Predict useful next tasks | Simulated user feedback / reward model | Synthetic + event-derived | Task prediction usefulness, acceptance proxy | GitHub |
| OmniMMI | OmniMMI | Streaming multimodal interaction | Continuous video, audio, speech, and dialogue history | Trigger proactive alerts and turn-taking while supporting reactive QA | Standing user queries | Human-annotated benchmark | Alert timing, turn-taking, streaming QA, multi-turn interaction | CVPR 2025 · Website |
| ESTP-Bench | Eyes Wide Open | Egocentric streaming video | Past/current first-person frames and evolving questions | Coherent just-in-time responses during perception | Evolving question context | Benchmark with instruction-tuning resources | ESTP-F1; coherence, response timing, efficiency | Paper · Code |
| StreamGaze | StreamGaze | Gaze-conditioned egocentric video | Causal video plus gaze trajectories | Anticipation and gaze-triggered alerts, alongside past/present reasoning | Measured gaze as an attention cue | 285 videos; 8,521 QA pairs across 10 tasks | Task-specific reasoning/alert evaluation; proactive tasks reported separately | Paper · Code |
| SOVBench-O / SOVBench-T | StreamOV | Streaming omni-video interaction | Continuous audio/video and multi-turn queries | Contextual inference; trigger on present evidence, stay silent for absent evidence | Standing queries and dialogue history | O: 172 sessions / 1,739 QA turns; T: 226 positive/negative samples | QA accuracy; trigger precision, recall, F1 | Paper / protocol |
| Live Gaming Benchmark | Proact-VL | Live companion commentary/guidance | Game video and scenario-conditioned interaction | Choose timely commentary, co-commentary, or guidance | Scenario roles / user guidance request | Game and guidance evaluation data | LLM Score, CC, F1, Time-Diff, PAUC | Code / protocol · Dataset |
| Full-Duplex Floor Selection | Full-Duplex Speech Models Take the Floor When Asked, Not When Needed | Always-on speech interaction | Context-matched monologues with controlled false-fact, hazard, missing-word, address, and silence triggers | Self-select to speak because intervention is needed, not merely because the floor is open | Speaker monologues rather than personalized users | Ten-condition controlled diagnostic | Speech-initiation probability, non-empty response rate, false-fact challenge, hazard warning | arXiv |
| Live Assistant | Live Assistant | Livestream social assistance | Native audio-video, comments, gifts, viewer dynamics, and room metadata | Choose silence, private memory, or a grounded message to the right recipient | Hosts and viewers reconstructed from real livestreams | 275 clips and 13,812 human-reviewed decision intervals | State, recipient, and task accuracy | arXiv |
| ProAction | Cognitive Action Reasoning for Proactive Robots | Human-centered embodied assistance | Visual observations, audio, and text without an action instruction | Infer a cognitively grounded high-level robot action from human and environmental cues | Human-refined action judgments informed by affective theory of mind | 10,000 samples across 12 scenarios and five scenes | Action reasoning, modality contribution, subject-disjoint transfer, human evaluation | arXiv |
| ContextAgentBench | ContextAgent | Wearable / open-world | Video, audio, notifications, persona context | Predict proactive services and tool calls | Persona-conditioned users | Benchmark | Service prediction, tool-call success | GitHub |
| FingerTip 20K | FingerTip 20K | Mobile | Android trajectories and user history | Suggest tasks and execute them personally | Long-term mobile users | Real trajectories | Task suggestion, personalized execution | GitHub |
| ProactiveEval | ProactiveEval | Proactive dialogue | Generated environments and targets | Plan targets and guide dialogue | Simulated users | Synthetic benchmark | Target planning, dialogue guidance, target density | GitHub |
| PROBE | Beyond Reactivity | Web / personal datastore | User priorities and personal documents | Discover and resolve bottlenecks | Priority profile | Synthetic personal datastore | Bottleneck discovery, action success | GitHub |
| UserVille | Training Proactive and Personalized LLM Agents | SWE and research tasks | Vague task prompts and user simulators | Ask useful questions and adapt to preferences | Preference-aware LLM users | Simulated environment | Productivity, proactivity, personalization | GitHub |
| ClarifyBench | Structured Uncertainty guided Clarification | Tool-using agents | Ambiguous tool calls and argument domains | Select the highest-value clarification and stop when further questions are not worth their cost | LLM-based interactive users | Dynamic multi-turn benchmark | Coverage, task success, question count, When2Call accuracy | Findings of ACL 2026 · arXiv |
| ChronosBench | Long-term Task-oriented Agent | Long-term task dialogue | Dynamic environment events and history | Maintain intent and follow up on triggers | User intent state over time | Synthetic benchmark | Intent-conditioned monitoring, event-triggered follow-up | arXiv |
| ProAgentBench | ProAgentBench | Real workflows | Workflow logs and long-term history | Decide when and how to assist | Real workflow users | Real data | Timing, content quality, long-history usefulness | arXiv |
| ProactiveMobile | ProactiveMobile | Mobile | Phone context, state, and API list | Infer latent intent and plan API sequence | Mobile context user | Synthetic / offline benchmark | API sequence success, proactive intelligence | CVPR 2026 · arXiv |
| ProEvent | ProEvent | Event tracking | Future events and reminders | Maintain future-event obligations | Event-driven user needs | Benchmark | Event tracking, reminder correctness | OpenReview |
| PIRA-Bench | PIRA-Bench | GUI | Continuous GUI screenshots | Recommend proactive intents | GUI user state | Benchmark | Intent recommendation accuracy, timing | Dataset |
| GUIDE | GUIDE | GUI / desktop workflows | Screen recordings with think-aloud narration | Detect behavior state, infer intent, and predict helpful assistance | Novice GUI users | Video benchmark | Behavior state detection, intent prediction, help prediction | Website · Dataset |
| ProactiveBench (MLLM) | ProactiveBench / Trento | Multimodal perception | Images with occlusion, poor quality, or ambiguity | Ask for help when visual evidence is insufficient | Visual collaborator | Repurposed visual datasets | Help-seeking, false positive rate, RL generalization | Dataset |
| Pare-Bench | Pare | Multi-app digital environment | FSM app states and active user simulation | Intervene, execute, or stay silent | Active simulated users | Simulator | Intervention timing, task success, user disruption | GitHub |
| KnowU-Bench | KnowU-Bench | Android personal agents | Behavior logs, app states, preferences | Clarify, act, personalize, and respect consent | Personalized Android users | Emulator benchmark | Task success, consent handling, rejection response | GitHub |
| LatentNeeds-Bench | PASK | Always-on personal assistance | 100 multi-turn sessions and 3,936 real-speech turns | Detect latent demand, choose fast or full assistance, or remain silent | User-consented speech transcripts | Human-refined benchmark | Demand, no-demand, balanced accuracy, latency | Website |
| CogEval-Bench | CogniFold | Proactive memory | Streaming events and concept graph | Surface emergent intents from memory structure | Implicit user memory graph | Benchmark | Concept emergence, cognitive structure, proactive surface | Dataset |
| MemEye | MemEye | Multimodal memory | Visual episodes and temporal state | Retrieve visual evidence and track changing state | Memory-dependent visual user | Diagnostic benchmark | Visual memory granularity, temporal reasoning | arXiv |
| ProCodeBench | Proactive Coding Assistants | IDE / coding | Real VS Code traces | Predict coding intent and assistant value | Real developers | Real traces | Sim-to-real gap, intent prediction, assistance quality | arXiv |
| ProactBench | Beyond What The User Asked For | Multi-turn dialogue | Incrementally disclosed persona-grounded details | Surface unstated needs at emergent, critical, and recovery triggers | Synthetic personas and user agent | 198 curated dialogues / about 624 triggers | Pass, partial, fail by trigger phase; human and judge agreement | arXiv |
| OmniPro | OmniPro | Omni-modal streaming video | Continuous audio-visual streams with multiple triggers | Decide when and what to report without polling | Standing natural-language instruction | 2,700 human-verified samples | Probe accuracy, online content/timing, over-trigger penalties, long-horizon retention | Website |
| EgoPro-Bench | EgoPro-Bench | Egocentric streaming assistance | Video stream, scene history, and simulated user profile | Trigger personalized assistance or remain silent | Simulated user memories and intentions | 2,400 evaluation videos plus 12,000 training videos | Precision, recall, F1, mIoU, trigger hit, response quality | arXiv |
| Claw-Anything | Claw-Anything | Always-on personal assistants | Months of activity, cross-service state, GUI and CLI across devices | Anticipate needs and act amid irrelevant or conflicting events | Simulated personas and event histories | 200 evaluation tasks | pass@1, context scaling, proactive-versus-reactive task gap | arXiv |
| π-Bench | π-Bench | Personal assistant workflows | Persistent workspaces, files, profiles, tasks | Resolve hidden intents in long-horizon workflows | Persona and workspace state | Benchmark | Proc, Comp, hidden-intent resolution | GitHub · Dataset |
| ProActEval | Anticipate and Learn | Proactive assistant | Dialogue history, persistent memory, idle-time context | Anticipate future needs and gather evidence | User profile and future need chain | Benchmark | User effort, hallucination reduction, proactive utility | GitHub |
| VitaBench 2.0 | VitaBench 2.0 | Long-term personalized interaction | Multi-session user interaction sequence | Extract, update, and use preferences; acquire missing info | Long-term user profile | Benchmark | Preference extraction, memory use, proactive acquisition | GitHub |
| Ψ-Bench | Ψ-Bench | Persuasive dialogue | User profiles and simulated client dialogues | Tailor influence strategies to personas | Profile-conditioned simulated clients | Benchmark | Persuasion quality, profile use, dialogue quality | GitHub |
| GuidanceSalesBench | See, Infer, Intervene | Smart retail | Pre-interaction videos, state manifests, candidate actions, and outcomes | Choose greet, elicit, inform, recommend, or hold | Scripted customer behavior and a staged pilot | Small human-checked benchmark | Action macro F1, state prediction, outcome simulation | arXiv |
| Int-Bench | AI Assistants Overassist | AI tutoring | Incremental student reasoning traces | Decide whether, when, and how much to intervene | LLM-simulated student plus human teacher comparison | 1,500 problems across three domains | Intervention frequency/timing, immediate helpfulness, transfer, content directness | arXiv |
| AskBench | When and What to Ask | Clarification dialogue | Intent-deficient questions and false-premise prompts | Ask a targeted question, correct a premise, or answer directly | Judge-loop simulated user | Interactive benchmark | Accuracy, rubric adherence, interaction efficiency | ACL |
| VibeLifeBench | VibeLifeBench | Long-horizon life assistance | Multi-week timelines, 22 mock services, silent world mutations | Decide when to act, ask, or stay silent while preserving a coherent plan | Persona, implicit constraints, authorization boundaries | Scripted living-world benchmark | Weighted stage checks, avg@3, max@3, interaction cost | arXiv |
| JarvisBench | JarvisBench | Human-agent attention coordination | Ongoing single-agent trajectories and coupled multi-agent workstreams | Recognize user-owned decisions, request judgment, and inject scoped guidance | Concise benchmark user decisions | Adapted public agent tasks | Task outcome gain, request count, attention efficiency, response quality, latency | Website · GitHub |
| Act2Intention Bench | Act2Intention | Mobile GUI | Continuous personalized intention-action trajectories | Predict and suggest the next intention, then execute after confirmation | 90 real users plus generated personas | Real + synthetic trajectories | Understanding accuracy, prediction accuracy, execution success rate | GitHub |
| StreamSoccer | StreamSoccer | Streaming soccer commentary | Causal match video, active event state, recent events, historical records | Select current, recent, or historical commentary, or remain silent | Broadcast audience rather than an individualized user | SoccerNet + MatchTime derived dataset | BLEU-4, CIDEr, BERTScore-F1, output coverage, real-time factor | arXiv |
| OmniAssistBench | OmniAssistBench | Streaming task assistance | Continuous video, user goal, interaction history, source-video priors | Guide the user, interpret visual prompts, and delay advice until the relevant event | Simulated interaction path reconstructed from source video | Expert-curated multi-turn video benchmark | 100-point assistant score; visual-prompt, history, and delayed-response diagnostics | Website · arXiv |
| Interactive Visual Grounding | When Seeing Is Not Enough | Interactive visual grounding | Image, incomplete target description, and follow-up dialogue | Ask questions, integrate answers, and identify the intended target | Human description providers and task-level human baselines | Four visual contexts × four interaction protocols | Grounding performance, task-level human gap, confidence calibration | arXiv |
| MMPCBench | MMPCBench | Multimodal critique | Text-image inputs containing one of 12 error subcategories | Detect, diagnose, and resolve flawed input without a checking prompt | User supplies a potentially faulty premise | Curated benchmark across four primary error types | Detection, diagnosis, resolution, reasoning–answer alignment | GitHub · arXiv |
| ProReady-QA | StreamReady | Long streaming-video QA | Continuous video, standing question, and dialogue history | Wait until answer evidence is sufficient, then answer without unnecessary delay | Question and annotated acceptable evidence windows | Five-task human-annotated benchmark | Accuracy, Answer Readiness Score, effective accuracy | CVPR 2026 · Website |
| PROS-Bench | Beyond Instruction-Driven Editing | Scientific-poster editing | Source paper plus editable PPTX poster | Discover source-grounded problems, request user acceptance, repair native objects, and validate outcomes | Explicit issue acceptance and final commit control | 120 papers and 320 editable posters, with a 120-poster matched primary core | Diagnosis quality, accepted-target resolution, outcome uplift, decline rate | arXiv |
| RPCBench | RPCBench | Recommendation critique | Evidence-grounded request with a corrupted premise | Detect, localize, and handle premise failures before recommending | User request plus visible profile, item, or constraint evidence | 4,623 instances across five domains and ten failure types | Detection, localization, handling strategy, evidence faithfulness | GitHub · arXiv |
| CONFLICTGUI | Do GUI Agents Know When Not to Act? | GUI execution safety | Instruction, GUI screenshot, action history, and feasibility evidence | Terminate instead of executing when the instruction or GUI state is conflicting | User-issued instruction that may contain a benign mistake | 2,364 feasible, 1,122 instruction-internal conflict, and 1,174 instruction-GUI conflict instances | Conflict-task success, termination behavior, feasible-task preservation | GitHub · arXiv |
| CC-Mediation | CC-Mediation | Cross-cultural conflict dialogue | Ten-turn conflict dialogue and intervention trajectory | Decide when to intervene and how to mediate for persistent stance improvement | Dialogue participants with culturally grounded value conflicts | 1,661 dialogues; 1,503 train and 158 evaluation | Trajectory AUC, signed Wasserstein-1 stance shift, timing and strategy failure | GitHub · arXiv |
| OR-Clarify | Ask Before You Optimize | Interactive optimization | Partial problem brief, dialogue history, and hidden formulation requirements | Ask formulation-critical questions, avoid silent assumptions, and stop when ready | Simulated user restricted to answering the current question | 100 cases and 178 hidden slots across P0/P1/P2 severity | Core/all-slot exact recovery, stopping behavior, silent assumptions, question count | arXiv |
| TIMELI | Time-Aware Assistive Navigation | Assistive outdoor navigation | Egocentric video, route plan, instruction history, and simulated user state | Issue concise safety guidance at the right time or remain silent | Blind-user needs informed by mobility guides and interviews | 67,000 synthetic navigation videos plus annotated real-world transfer videos | Timing F1/AUC, BLEU-4, ROUGE-L, conciseness, Navigation Quality Score, collision and instruction rates | Website · arXiv |
| IdeaAMBIG | IdeaAMBIG | Research coding / specification | Method specification grounded in papers, code, issue threads, and reproduction artifacts | Detect implementation-critical gaps and generate clarification actions instead of assuming | A competent implementer or coding agent as the downstream consumer | 660 instances: 163 real gaps and 497 controlled synthetic gaps | Codification-readiness accuracy, Macro Defect Recovery Rate, clarification-action success | arXiv |
| Physical Experiment Selection | New Evidence, Same Choice | Active physical reasoning | One measurement image, a question, possible worlds, experiment choices, and costs | Stop and answer when evidence suffices or select the cheapest resolving experiment | Task-defined evidence requirement rather than a conversational user | 144 parameter families and 576 core decisions across sliding, bouncing, and spring systems | Minimum-cost choice, both-correct matched-pair score, answer accuracy, perception and numerical controls | arXiv |
| ProMediConv | ProMediConv | Multi-party legal mediation | Dialogue history, mediation stage, party behavior states, and strategy choices | Proactively select a mediation strategy and improve party behavior through the dialogue | Simulated disputing parties reconstructed from complete legal cases | 972 cases with utterance-level labels for 11 strategies and four behavior states | Mean Attribute Difference, average turns, success and soft-success rates, strategy selection | GitHub · arXiv |
| PASTABench | PASTABench | Sequential agent safety | Multi-step tool-agent trajectories with accumulating risk evidence | Decide whether and when to interrupt inside the optimal intervention window | Risk annotations rather than a personalized user model | 1,139 trajectories across five risk categories and 13 subcategories | Optimal-timing intervention, earliest-signal and trigger turns, risk classification, lexical-robustness controls | arXiv |
| TWIST | TWIST | Conversational memory governance | Conversation history, belief changes, drafts, and sensitive-memory cases | Detect or block when memory warrants intervention while avoiding matched false alarms | Human-validated belief and draft-alignment judgments | Four proposed tracks; 161-item human-validated Track B v1.0 | Contradiction recall, hard-negative specificity, attribution, sensitive-recall governance | arXiv |
| TRACE | TRACE | Streaming video understanding | Causal video stream, audited evidence windows, standing instruction, and recorded system events | Trigger a response only when evidence is valid while controlling delay and workload | Instruction-defined trigger rather than a personalized user | 1,240 records from 517 videos; eight systems in eight configurations | Answer quality, delay, false alarms, missed windows, workload, completion, reliability | GitHub · arXiv |
| IntentFlux | When Users Change Their Minds | Multi-turn agent tasks | Executable task dialogue with superseded and withdrawn requirements | Recover the active intent and prevent stale instructions from affecting the final answer or tool action | Controlled user intent changes | Executable benchmark with a 627-case calibration and preserved task graders | Task score, fully correct rate, sensitivity to superseded information, state-recovery gain | arXiv |
| RobotEQ 3.0 | RobotEQ 3.0 | Personalized embodied social intelligence | Social context, user profile, and candidate robot actions | Predict the proactive action preferred by a specific user | Participant questionnaires and individual action preferences | Human-labeled personalized preference dataset | Personalized action-prediction accuracy, inter-annotator variance, profile contribution | arXiv |
| PROACTIVITY-GYM | Foundations of Proactive Agents | Multi-day proactive assistance | Stateful environment, resource schedule, persona, and intervention history | Jointly optimize useful work, compute timing, and evolving user trust | Persona-conditioned simulated users plus a 30-participant study | Simulation testbed across 23 model-harness configurations | Task capability, temporal allocation, trust, intervention misalignment | arXiv |
| APM-Bench | APM-Bench | Cross-session egocentric assistance | Intermittent egocentric video sessions with finite persistent memory | Retain and inject evidence for later responses and acknowledge when evidence is missing | Life-trajectory questions rather than an individualized preference model | 549 sessions, 104 trajectories, and 2,719 candidates | Utility, latency, storage, long-term recall, proactive assistance, missing-evidence awareness | arXiv |
| Drift-Bench++ | Beyond Oracle Communication | Interactive intent alignment | Executable tasks with miscommunication, finite patience, and silent intent shifts | Ask effectively and continuously track the user's current intent | Diverse simulated users plus deployed-session validation | Controlled benchmark-construction pipeline with executable graders | Task grounding, user realism, inquiry effectiveness, intent-shift adaptation | arXiv |
| DuplexAct-Bench | DuplexAct-Bench | Full-duplex speech interaction | Bilingual streaming speech across pre-session, in-session, and no-explicit conditions | Interrupt, yield, initiate, stay silent, or backchannel at the right time | Scripted English and Chinese interaction conditions | 1,290 trials across six behaviors; 12 evaluated systems | Timing and content by behavior and condition | Website · arXiv |
| SWE-Intervene | Learning When and How to Intervene | Coding-agent execution | Repository trajectory, pre-action context, proposed action, and outcome-derived label | Allow, autonomously redirect, or pause for human assistance before execution | Human assistance is an escalation action rather than a live user model | Action-level software-engineering trajectories | Task completion gain, intervention choice, actionable feedback, token consumption | arXiv |
| OverAct | OverAct | Tool-call privacy and authorization | User request, private service APIs, and candidate tool calls | Prevent access beyond the minimum scope justified by the request | Explicit request as the authorization boundary | Controlled benchmark across eight privacy-sensitive domains | Deterministic excess-access score, privacy-oriented excess, task preservation | arXiv |
Selection Guide
| Research Question | Start With | Why |
|---|---|---|
| When should a proactive agent interrupt? | DuplexAct-Bench, Int-Bench, JarvisBench, Live Assistant, PASTABench, PROACTIVITY-GYM, RealHumanEval, Pare-Bench, ProAgentBench | They expose attention needs, active silence, acceptance, disruption, safety windows, or trust effects rather than only final task success. |
| How do we evaluate hidden or changing intent? | Act2Intention Bench, Drift-Bench++, IntentFlux, π-Bench, GUIDE, PIRA-Bench, ProactiveMobile, ProactiveBench | They require agents to infer goals that are incomplete, miscommunicated, or superseded over time. |
| How do we evaluate long-term intent maintenance? | VibeLifeBench, APM-Bench, ChronosBench, Drift-Bench++, VitaBench 2.0, π-Bench | They require persistent state across time or sessions; VibeLifeBench also advances the world while the agent is not being prompted. |
| How do we evaluate personalization? | RobotEQ 3.0, KnowU-Bench, FingerTip 20K, VitaBench 2.0, UserVille, Ψ-Bench | They include profiles, preferences, or user-specific trajectories. |
| How do we evaluate memory as a proactive substrate? | APM-Bench, CogEval-Bench, MemEye, TWIST, VitaBench 2.0, ProActEval | They test memory formation, retrieval, evidence preparation, intent emergence, missing-evidence awareness, or whether memory should trigger an intervention. |
| How do we evaluate computer-use agents? | Act2Intention Bench, CONFLICTGUI, SWE-Intervene, OverAct, ProAgentBench, GUIDE, PIRA-Bench, KnowU-Bench, Pare-Bench, ProCodeBench | They connect proactive behavior to GUI, mobile, IDE, tool calls, or workflow contexts; CONFLICTGUI and OverAct test restraint, while SWE-Intervene tests pre-execution control. |
| How do we evaluate proactive streaming video or speech? | ESTP-Bench, StreamGaze, SOVBench, Live Gaming Benchmark, APM-Bench, DuplexAct-Bench, Live Assistant, TRACE, TIMELI, ProReady-QA, OmniAssistBench, OmniPro, EgoPro-Bench, OmniMMI, StreamArena, IPIBench, EgoServe, StreamSoccer | They require causal processing of an ongoing stream and timely outputs rather than offline clip answering; TRACE audits execution conditions and DuplexAct-Bench adds proactive initiation and active silence. |
| How do we evaluate proactive critique? | RPCBench, MMPCBench, AskBench | They require the assistant to challenge or repair a faulty premise instead of complying silently; RPCBench additionally tests localization, handling strategy, and evidence faithfulness in recommendation. |
| How do we evaluate ask-versus-stop clarification? | FinInteract, IdeaAMBIG, OR-Clarify, ClarifyBench, AskBench | They evaluate whether missing information justifies another question and whether the agent can stop without silently assuming critical information; FinInteract additionally tests whether the answer is integrated under competing valid interpretations. |
| How do we evaluate social intervention outcomes? | ProMediConv, CC-Mediation, ProMediate, ProACT Collaboration Bench | They evaluate intervention strategies or points and downstream collaborative or interpersonal effects rather than response quality alone. |
| How do we evaluate proactive evidence acquisition? | Physical Experiment Selection, Value of Information, Interactive Visual Grounding | They test whether an agent should answer now, ask for information, or acquire another observation, with explicit evidence sufficiency or cost. |