OpenVLHarness
Real-World Visual Agentic Tool-Use with Multimodal Memory
1University of California, Los Angeles · 2Apple
*Equal contribution
Architecture
Recorded trajectories, replayed step by step. Select an example, pause or resume playback, step through it, or click a panel to expand it.
User input
MLLM orchestrator
Environment Update
Tool call
Processed tool outputs
Interactive demo
The agent loop and perception tools run on our server; your model (OpenAI, or an OpenAI-compatible endpoint with tool calling) orchestrates, writes the coding tool's Python and summarizes web results.
- The trajectory appears here while the agent runs.
Per-benchmark results
Bold = best within each backbone group; underline = best across all shown backbones. Qwen3-VL results are means of three runs; GPT-6, GPT-5 and Kimi K3 results are single runs. GPT-5 and Kimi K3 are from the appendix (Table 6).
Analysis
Per-benchmark gains
Base and OpenVLHarness scores on each of the 23 benchmarks, per backbone. Average improvement over the base model ranges from +9.2 to +12.6 points across six backbones; with GPT-6 Sol, all 23 benchmarks improve.
Comparison with other harnesses
Average over 23 benchmarks with the same backbone. GPT-6 Luna: Base 43.2, Visual Sketchpad 42.8, Codex 41.5, OpenVLHarness 55.8. Qwen2.5-VL-7B: AdaReasoner 27.5, PixelReasoner 29.3, DeepEyesV2 26.7 (fine-tuned for tool use); OpenVLHarness 35.6 (original weights).
Component-wise ablation
Components added cumulatively to the no-tool baseline, Qwen3-VL-8B and 32B. Each cell shows the score and its change from the row above.
Shading: change from the previous configuration ( gain, loss). Source: paper Table 2.
Tool usage by domain
Share of tool calls in each domain, per backbone, over all evaluated questions. Micro counts every call; per-instance counts each tool at most once per question. Totals are given below each panel.
Source: paper Fig. 7 (appendix) telemetry. Micro view reproduces the figure.
Harness optimization
Kimi K3 critiques 9,980 failed Qwen3-VL-8B trajectories and assigns each a root cause and a remedy channel: new task-specific tools (A–C), harness text (Y), or none (Z). The dominant channel selects the workflow. Harness search edits the system prompt, tool descriptions, and output templates; tool generation composes existing capabilities into new tools. Both are run by a coding agent with model weights fixed.
Citation
@article{openvlharness2026,
title = {OpenVLHarness: Real World Visual Agentic Tool-Use with Multimodal Memory},
author = {Zhou, Yu and Yang, Cheng-Fu and Sun, Rui and Wu, Di and Peng, Sibo and Zhang, Mingyang and Yang, Zhen and Zhang, Ruiyi and Peng, Nanyun and Gan, Zhe and Chang, Kai-Wei},
journal = {arXiv preprint arXiv:XXXX.XXXXX},
year = {2026}
}