IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

Jinzhao Li1,2, Yinuo Chen1, Wenxuan Song1, Yijia Lei1, Yichi Zhang1, Honglei Yan2, Panwang Pan2,*, Miao Liu1,†

1College of AI, Tsinghua University  ·  2ByteDance  ·  Corresponding author  ·  *Project lead

Paper Code Dataset Coming soon Demo
1,831continuous videos
3,738interactive QA instances
3core task families
Demo

See IPIBench and IPI-Agent in motion.

A 3-minute walkthrough of our IPIBench and our IPI-Agent framework.

From reactive to proactive

From answering to proactively responding.

Real assistants must reason continuously, support multi-turn interaction, act at the right moment, and adapt as people change their minds.

Recent MLLMs excel at reactive question answering, yet real-world streaming assistants need to proactively reason over continuous visual inputs, respond at the right moment, and coordinate with users across multi-turn interactions. IPIBench evaluates this shift in dynamic video settings where users can add, modify, or cancel proactive requests while interleaving reactive questions.

IPIBench benchmark

We introduce the first benchmark for evaluating interactive proactive intelligence of MLLMs under streaming video settings, covering proactive monitoring, proactive task management, and interleaved reactive–proactive requests.

Systematic evaluation

We evaluate representative proprietary, open-source, and online streaming models, revealing unstable proactive triggering and weak multi-turn interaction coordination.

IPI-Agent framework

We propose a training-free agentic framework with an interaction-control policy and temporal-gating mechanism to improve triggering stability and multi-turn coordination.

IPIBench

Evaluating Interactive Proactive Intelligence of MLLMs

Three task families capture the shift from one-shot alerts to stateful, mixed-initiative interaction.

Overview of IPIBench task families and interaction patterns
A benchmark for real streaming behaviorSingle-turn tasks test timing, understanding, and repeated triggering. Multi-turn tasks test cancellation, modification, multi-task management, and reactive–proactive coordination.
Dataset at a glance

Diverse task types. Broad video sources.

IPIBench spans 1,831 videos and 3,738 QA instances across rich task types and broad egocentric/exocentric sources.

Task distribution

Share of all QA instances

3,738QA instances
Proactive Monitoring46.71%
Task Management25.07%
Interleaved Requests28.22%

Video duration

Durations range from short clips measured in seconds to long videos exceeding five minutes.

0–50 s55.74%
50–100 s30.45%
100–150 s8.54%
150–200 s2.31%
200–250 s0.92%
250–300 s0.68%
≥ 300 s1.36%

Video sources

Complementary first- and third-person viewpoints

42.65%57.35%
Egocentric781
Exocentric1,050
394Daily indoor
254Cooking
201Vlogs
170Handicraft
162Household chores
160Failure events
140Tutorials
115Sports
108Social & leisure
87Driving
40Movies
IPI-Agent

Always attentive, never intrusive.

A training-free agentic framework that turns existing offline MLLMs into more stable, stateful streaming assistants.

01
Interaction-control policy

Routes reactive queries and manages add, edit, and cancel instructions with explicit memory.

02
Temporal-gating mechanism

Combines recent frames with proactive memory to decide when to respond and when to stay silent.

03
Tool-based modularity

Separates gating, memory, and response generation for plug-and-play use across base MLLMs.

IPI-Agent architecture with interaction control, temporal gating, memory and response tools
Leaderboard

Current models are still far from human.

Scores (%) across all nine IPIBench sub-categories and three task-family averages.

best non-human score
ModelProactive MonitoringProactive Task ManagementInterleaved Reactive–Proactive
TimingUnder.Repeat.Avg.CancelModifyMultiAvg.R2PRuPRaPAvg.
Human Level98.0094.0092.0094.6798.0092.0098.0096.0096.0096.0098.0096.67
Proprietary Models · Offline
Gemini 3 Pro43.4924.9018.4028.9319.4421.8419.3920.2224.7617.6826.6923.04
Gemini 2.5 Pro45.3625.2613.2127.9424.5420.8116.7820.7121.1717.5428.2422.32
GPT-5.454.6423.528.9629.0438.4324.507.5723.5027.0415.0629.2823.79
GPT-4o50.5925.9913.6830.0911.5731.8814.1819.2127.0417.6826.1723.63
Open-source Models · Offline
LLaVA-OneVision-7B12.148.660.006.931.391.010.470.963.910.003.892.60
InternVL3-8B32.5122.313.7719.536.946.043.555.517.170.009.335.50
Qwen3-VL-8B43.1224.514.2523.961.3917.459.229.3524.4319.7523.0522.41
Qwen3.5-Plus35.2524.357.8922.5012.9631.5410.8718.4632.2522.0130.5728.28
GLM-4.6V50.4327.5910.8529.6215.2824.169.6916.3822.156.7739.3722.76
Open-source Models · Online
VideoLLM-online-8B14.630.141.825.537.871.686.625.390.001.101.550.88
Dispider18.720.651.827.06
Flash-VStream-7B5.910.000.001.972.310.000.000.770.000.000.780.26
Higher is better. Under. = Understanding; Repeat. = Repeated Proactiveness.Scroll horizontally to view all metrics →
IPI-Agent results

One framework, consistent gains.

IPI-Agent improves four representative MLLMs without additional training. Green values show absolute gains over each base model.

Gemini 3 Pro+18.50Task Management Avg.
GPT-5.4+5.87Task Management Avg.
Qwen3.5-Plus+18.72Task Management Avg.
Qwen3-VL-8B+17.39Task Management Avg.
IPI-Agent BaseProactive MonitoringProactive Task ManagementInterleaved Reactive–Proactive
TimingUnder.Repeat.Avg.CancelModifyMultiAvg.R2PRuPRaPAvg.
Gemini 3 Pro56.27+12.7835.67+10.7728.30+9.9040.08+11.1551.85+32.4135.23+13.3929.08+9.6938.72+18.5030.62+5.8622.38+4.7029.53+2.8427.51+4.47
GPT-5.457.20+2.5624.30+0.7810.85+1.8930.78+1.7448.61+10.1824.83+0.3314.66+7.0929.37+5.8728.01+0.9719.34+4.2833.42+4.1426.92+3.13
Qwen3.5-Plus52.80+17.5533.18+8.8318.42+10.5334.80+12.3052.78+39.8233.22+1.6825.53+14.6637.18+18.7232.25+0.0022.63+0.6235.95+5.3830.28+2.00
Qwen3-VL-8B46.62+3.5025.37+0.8615.09+10.8429.03+5.0744.91+43.5223.49+6.0411.82+2.6026.74+17.3927.36+2.9323.62+3.8730.57+7.5227.18+4.77
Score on top; absolute gain over the original base MLLM below.No additional training required.
Ablation study

Both components matter.

Removing either interaction control or temporal gating degrades performance, using Qwen3-VL-8B as the base model.

Variant / Δ from fullProactive MonitoringProactive Task ManagementInterleaved Reactive–Proactive
TimingUnder.Repeat.Avg.CancelModifyMultiAvg.R2PRuPRaPAvg.
Full IPI-Agent46.6225.3715.0929.0344.9123.4911.8226.7427.3623.6230.5727.18
w/o Interaction Control0.000.000.000.00−40.95−4.70−0.24−15.54−1.81−0.55−3.11−1.82
w/o Temporal Gating−3.50−0.86−10.84−5.07−6.95−5.37−1.65−4.66−0.32−2.21−3.63−2.05
Citation

Build on IPIBench.

If you find this work useful, please cite our paper.

@article{li2026ipibench,
  title={IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams},
  author={Li, Jinzhao and Chen, Yinuo and Song, Wenxuan and Lei, Yijia and Zhang, Yichi and Yan, Honglei and Pan, Panwang and Liu, Miao},
  journal={arXiv preprint arXiv:2605.27074},
  year={2026}
}