arxiv_cv 95% Match Research Paper AI Researchers (Multimodal),ML Engineers,Benchmark Developers 3 days ago

TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning

large-language-models › multimodal-llms

📄 Abstract

Abstract: The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as thinking-\textit{with}-images in chain-of-thought. Yet existing benchmarks fail to fully capture this advanced capability. Even Visual Search, the most common benchmark for current thinking-\textit{with}-images methods, tests only basic operations such as localization and cropping, offering little insight into more complex, dynamic, and tool-dependent reasoning. We introduce \textbf{TIR-Bench}, a comprehensive benchmark for evaluating agentic thinking-with-images across 13 diverse tasks, each requiring novel tool use for image processing and manipulation in chain-of-thought. We evaluate 22 multimodal large language models (MLLMs), from leading open-sourced and proprietary models to those with explicit tool-use augmentation. Results show that TIR-Bench is universally challenging, and strong performance requires genuine thinking-with-images capabilities. Finally, we present a pilot study comparing direct versus agentic fine-tuning.

Authors (9)

Ming Li

Jike Zhong

Shitian Zhao

Haoquan Zhang

Shaoheng Lin

Yuxiang Lai

+3 more

Submitted

November 3, 2025

arXiv Category

cs.CV

arXiv PDF

Key Contributions

Introduces TIR-Bench, a comprehensive benchmark for evaluating agentic thinking-with-images reasoning across 13 diverse tasks requiring novel tool use for image processing and manipulation. It evaluates 22 MLLMs, highlighting the universal challenge of these tasks and the need for advanced capabilities beyond basic operations.

Business Value

Provides a standardized and challenging evaluation framework for multimodal AI systems, enabling better development and comparison of models for applications requiring sophisticated visual understanding and manipulation, such as automated image editing or robotic vision.

Paper Metadata

Innovation Type

Benchmark Creation

Deployment Feasibility

High for using the benchmark; moderate for developing models that perform well on it.

Limitations Addressed

Existing benchmarks fail to fully capture advanced capabilities of models like OpenAI o3, which can intelligently create and operate tools for image problem-solving (thinking-with-images). Visual Search benchmarks are too basic.

Technical Tags

Multimodal LLMsVisual ReasoningAgentic AITool UseChain-of-ThoughtImage ManipulationBenchmarkEvaluation

Research Topics

Multimodal AIVisual ReasoningAgentic SystemsAI BenchmarkingLarge Language Models

Methods & Architectures

Chain-of-Thought ReasoningTool Use IntegrationImage Processing OperationsBenchmark Design Multimodal Large Language Models (MLLMs)Agentic Models

Applications & Tasks

Image Analysis Content Creation Robotics Human-Computer Interaction Complex Visual ReasoningDynamic Image ManipulationTool-dependent Problem Solving Evaluating agentic thinking-with-images capabilitiesAssessing multimodal LLM performance on complex visual tasksBenchmarking tool use in image processing

Datasets & Benchmarks

Benchmarks

TIR-Bench (13 tasks)

Related Fields

Computer VisionNatural Language ProcessingArtificial IntelligenceRobotics

Keywords

Multimodal LLMVisual ReasoningAgentic AITool UseChain-of-ThoughtImage ManipulationBenchmarkEvaluationMLLMOpenAI o3TIR-BenchProblem Solving

Academic Context

#Multimodal AI#Visual Reasoning#Agentic Systems#AI Benchmarking#Large Language Models

Companies & Organizations

Companies Mentioned

OpenAI

Commercial Potential

Potential Products

Advanced Image Editing SoftwareAI Assistants for Creative TasksRobotic Vision Systems

Target Industries

Media and EntertainmentE-commerceRoboticsSoftware Development

Use Case Examples

Automated photo retouching based on complex instructionsAI agents that can modify images to achieve specific aesthetic goalsRobots that use visual feedback to manipulate objects

Competitive Edge

Establishes a new, more rigorous benchmark for evaluating advanced multimodal reasoning and tool-use capabilities, pushing the state-of-the-art beyond simpler visual tasks.

Market Opportunity

Large, driven by the rapid advancement and commercialization of multimodal AI.

Revenue Models

Benchmark usage feesconsulting services for model evaluation.

Resource Requirements

Compute Needs

Significant for training and evaluating MLLMs on the benchmark.

Data Requirements

The benchmark itself comprises tasks and potentially associated data.

Deployment Constraints

Complexity of models required to perform well on the benchmark.

Production Readiness

Maturity Level

Benchmark Release

Time to Market

N/A (benchmark release)

Patent Potential

Low, focused on evaluation methodology.

View Full Paper Back to Papers