Must Read
A layered starting path, selected for coverage of different mechanisms and evidence, rather than publication date or a shared leaderboard. Read Tier 1 first, choose an application or mechanism in Tier 2, then use Tier 3 to evaluate it. Tiers are a reading order, not a quality ranking. Products and code are included alongside papers.
| Route | Representative resources | Why start here |
|---|---|---|
| Tier 1 · Foundations — define proactivity, build a baseline, understand interruption | ||
| Human-centered definition | Towards Human-centered Proactive Conversational Agents SIGIR 2024 · conceptual foundation | Intelligence, adaptivity, and civility provide a starting vocabulary for useful, respectful initiative. |
| Canonical agent baseline | Proactive Agent · Code · Notes ICLR 2025 · method + benchmark | Connect desktop event streams to task anticipation and ProactiveBench evaluation. |
| User experience and timing | Need Help? · Notes CHI 2025 · human factors | Understand when proactive IDE help is useful and how users retain control. |
| Tier 2 · Mechanisms and systems — choose the kind of initiative you want to build | ||
| Sensory context → assistance | ContextAgent · Code NeurIPS 2025 · context-aware agent | Move from desktop logs to open-world sensory context and tool-mediated help. |
| Persistent agent activation | OpenClaw · Heartbeat docs Proactivity SDK Runtime projects · see Projects & Products | Compare periodic attention checks with durable goals and model-selected wake cadence. |
| Product-level initiative | OpenAI dot MineContext Always-on product / ambient desktop project | Compare background research and connected-app context with proactive desktop summaries and tips. |
| Latent needs and persistent memory | PASK / IntentFlow · Code / project Notes Technical report · need detection and assistance routing | Connect silent / fast / full assistance decisions to ongoing conversations and hierarchical personal memory. |
| Personal routines across devices | PersonalAlign / HIM-Agent Personal Agents guide Mobile user history · preference and routine memory | Separate completing omitted preferences from anticipating an unrequested routine; personal assistance extends beyond desktop interfaces. |
| Wearable and physical-world help | Satori ProMemAssist AR guidance · working-memory-aware intervention | Ground assistance in inferred user/task state and weigh the benefit of help against interruption. |
| Embodied ask / act decisions | PACT PACE Continual robot assistance · action-completion timing | Compare asking versus acting from interaction history with timing physical help from observed task progress. |
| Visual events → speak / silence | JoyAI-VL-Interaction · Code MOSS-VL-Realtime · Code Streaming interaction models | Study internally controlled response timing and vision-triggered intervention, including JoyAI's delegation decision. |
| Evidence and memory → timely answer | OneStreamer · Code · Notes StreamReady Streaming memory and answer readiness | Contrast proactive evidence recording with waiting until a standing question has enough evidence. |
| Full-duplex interaction → agent work | MiniCPM-o 4.5 · Code Gander · Code Simultaneous perception, speech, and asynchronous work | Study speech/listening control, scene-driven comments, and the bridge to a background reasoning agent. |
| Tier 3 · Evaluation and boundaries — test the route you chose | ||
| Real workflows and personalization | ProAgentBench KnowU-Bench · Code | Pair real desktop traces with personalized, consent-aware mobile evaluation. |
| GUI intent recommendation | PIRA-Bench / PIRF · Project Data Benchmark with a memory-aware recommendation baseline | Evaluate latent future intents, user profiles, interleaved tasks, and rejection of noisy visual context before instruction-following execution. |
| Long-horizon initiative | π-Bench · Code VibeLifeBench | Evaluate hidden intents and act / ask / stay-silent decisions over evolving personal workflows. |
| Streaming timing and intent | OmniMMI StreamGaze · Code Streaming evaluation routes | Compare multimodal alert/turn timing with gaze-conditioned intent; inspect proactive subsets separately from general streaming QA. |
| Personalized physical-world assistance | EgoPro-Bench Vinci2 / EgoServe Egocentric timing, user context, and silence | Connect streaming to personalized assistance while keeping recorded-video evaluation separate from live wearable deployment. |
| Silence and oversight | Why2Speak · Notes Benchmark Matrix | Audit act-versus-abstain decisions, class imbalance, and whether exposed reasoning changes the policy being inspected. |