OpenVLHarness

Real-World Visual Agentic Tool-Use with Multimodal Memory

Yu Zhou1*, Cheng-Fu Yang1*, Rui Sun1*, Di Wu1*, Sibo Peng1, Mingyang Zhang1, Zhen Yang2, Ruiyi Zhang2, Nanyun Peng1, Zhe Gan2, Kai-Wei Chang1

1University of California, Los Angeles  ·  2Apple
*Equal contribution

Gains across model families Average over 23 benchmarks · Base vs. + OpenVLHarness
Base MLLM + OpenVLHarness
GPT-6: performance vs. cost 12 benchmarks · labels: reasoning effort

Architecture

Recorded trajectories, replayed step by step. Select an example, pause or resume playback, step through it, or click a panel to expand it.

Domain
User input
Answer: …
MLLM orchestrator
idle
Environment Update
Tool call
Processed tool outputs
Capability Layer
Perception
Zoom-in
Grounding
OCR
Depth
Camera pose
Search
Text search
Image search
Web visit
Coding
Coding Agent
Models & services
Segment Anything Depth Anything PaddleOCR‑VL Search engines Coding modelorchestrator MLLM Image processors

Interactive demo

The agent loop and perception tools run on our server; your model (OpenAI, or an OpenAI-compatible endpoint with tool calling) orchestrates, writes the coding tool's Python and summarizes web results.

Connecting to the demo server…
Samples web = requires a Serper key
Orchestrator model

Keys are used for the request only and are not stored.

Trajectory
  1. The trajectory appears here while the agent runs.

Per-benchmark results

Bold = best within each backbone group; underline = best across all shown backbones. Qwen3-VL results are means of three runs; GPT-6, GPT-5 and Kimi K3 results are single runs. GPT-5 and Kimi K3 are from the appendix (Table 6).

Analysis

Per-benchmark gains

Base and OpenVLHarness scores on each of the 23 benchmarks, per backbone. Average improvement over the base model ranges from +9.2 to +12.6 points across six backbones; with GPT-6 Sol, all 23 benchmarks improve.

Base + OpenVLHarness

Comparison with other harnesses

Average over 23 benchmarks with the same backbone. GPT-6 Luna: Base 43.2, Visual Sketchpad 42.8, Codex 41.5, OpenVLHarness 55.8. Qwen2.5-VL-7B: AdaReasoner 27.5, PixelReasoner 29.3, DeepEyesV2 26.7 (fine-tuned for tool use); OpenVLHarness 35.6 (original weights).

Component-wise ablation

Components added cumulatively to the no-tool baseline, Qwen3-VL-8B and 32B. Each cell shows the score and its change from the row above.

Shading: change from the previous configuration ( gain, loss). Source: paper Table 2.

Tool usage by domain

Share of tool calls in each domain, per backbone, over all evaluated questions. Micro counts every call; per-instance counts each tool at most once per question. Totals are given below each panel.

Source: paper Fig. 7 (appendix) telemetry. Micro view reproduces the figure.

Harness optimization

Kimi K3 critiques 9,980 failed Qwen3-VL-8B trajectories and assigns each a root cause and a remedy channel: new task-specific tools (A–C), harness text (Y), or none (Z). The dominant channel selects the workflow. Harness search edits the system prompt, tool descriptions, and output templates; tool generation composes existing capabilities into new tools. Both are run by a coding agent with model weights fixed.

Failure causes and remedy channels by domain Qwen3-VL-8B · 9,980 failures · unweighted mean over datasets per domain
Remedy profiles and specialization gains Qwen3-VL-8B · gains over the unoptimized harness
New task-specific tools (A+B+C) Harness text (Y) No remedy (Z)

Citation

@article{openvlharness2026,
  title   = {OpenVLHarness: Real World Visual Agentic Tool-Use with Multimodal Memory},
  author  = {Zhou, Yu and Yang, Cheng-Fu and Sun, Rui and Wu, Di and Peng, Sibo and Zhang, Mingyang and Yang, Zhen and Zhang, Ruiyi and Peng, Nanyun and Gan, Zhe and Chang, Kai-Wei},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}