Skip to content

AI Models with Vision

236 multimodal AI models that support image input and visual understanding, ranked by composite score. These models can analyze images, screenshots, diagrams, and documents alongside text prompts.

Vision Models

236

Of All Models

57%

Providers

27

Avg Score

62

Vision Model Rankings

All 236 models with vision capability, sorted by composite score. Score is computed from capabilities, pricing, context window, recency, output capacity, and versatility.

#ModelScore
1Claude Opus 4.7 (Fast)Anthropic90.9
2GPT-5.5 ProOpenAI90.9
3GPT-5.4 ProOpenAI90.9
4GPT-5.5 Pro (batch)OpenAI88.4
5GPT-5.4 Pro (batch)OpenAI88.4
6GPT-5.2 ProOpenAI88.1
7GPT-5 ProOpenAI86.4
8GPT-5.2 Pro (batch)OpenAI84.1
9Claude Opus 5 (Fast)Anthropic78.4
10Claude Fable Latest~anthropic78.4
11Claude Fable 5Anthropic78.4
12Claude Opus 4.8 (Fast)Anthropic78.4
13o3 ProOpenAI77.3
14GPT-5 Pro (batch)OpenAI76.4
15o1-proOpenAI76.4
16o1-pro (batch)OpenAI76.4
17GPT-5.6 Sol ProOpenAI73.4
18GPT-5.6 SolOpenAI73.4
19OpenAI GPT Latest~openai73.4
20GPT-5.5OpenAI73.4
21Claude Opus 4.1Anthropic73.1
22Claude Opus 5Anthropic72.1
23Claude Fable 5 (batch)Anthropic72.1
24Claude Opus 4.8Anthropic72.1
25Claude Opus Latest~anthropic72.1
26Claude Opus 4.7Anthropic72.1
27Claude Opus 4.6Anthropic71.9
28Google Gemini Pro Latest~google71.8
29Gemini 3.1 Pro Preview Custom ToolsGoogle71.8
30Gemini 3.1 Pro PreviewGoogle71.8
31Fugu Ultrasakana71.7
32Claude Opus 4Anthropic71.1
33Gemini 3.5 FlashGoogle71
34Gemini 3.6 FlashGoogle70.6
35Google Gemini Flash Latest~google70.6
36Gemini 3.1 Pro Preview (batch)Google70.3
37Gemini 3.5 Flash (batch)Google69.9
38GPT-5.4 Image 2OpenAI69.9
39Gemini 3.6 Flash (batch)Google69.7
40GPT-5.6 Sol Pro (batch)OpenAI69.7
41GPT-5.6 Sol (batch)OpenAI69.7
42GPT-5.5 (batch)OpenAI69.7
43GPT-5.4OpenAI69.7
44Claude Sonnet 4.6Anthropic69.6
45Gemini 3.5 Flash LiteGoogle69.4
46Nano Banana Pro (Gemini 3 Pro Image)Google69.3
47Gemini 3.5 Flash Lite (batch)Google69.1
48Gemini 3.1 Flash LiteGoogle69.1
49Gemini 3.1 Flash Lite PreviewGoogle69.1
50Claude Opus 5 (batch)Anthropic69
51Claude Opus 4.8 (batch)Anthropic69
52Claude Opus 4.7 (batch)Anthropic69
53Gemini 3.1 Flash Lite (batch)Google68.9
54GPT Chat LatestOpenAI68.8
55Claude Opus 4.6 (batch)Anthropic68.7
56Claude Sonnet 5Anthropic68.4
57Anthropic Claude Sonnet Latest~anthropic68.4
58GPT-5.3-CodexOpenAI68.4
59Gemini 3 Flash PreviewGoogle67.9
60GPT-5.4 (batch)OpenAI67.8
61Claude Sonnet 4.6 (batch)Anthropic67.7
62Claude Opus 4.5Anthropic67.7
63Gemini 3 Flash Preview (batch)Google67.6
64o1OpenAI67.5
65GPT-5.6 Terra ProOpenAI67.4
66GPT-5.6 Terra Pro (batch)OpenAI67.4
67GPT-5.6 TerraOpenAI67.4
68GPT-5.6 Terra (batch)OpenAI67.4
69MoonshotAI Kimi Latest~moonshotai67.4
70Gemma 4 26B A4B (free)Google67.4
71Gemma 4 31B (free)Google67.4
72Claude Opus 4.1 (batch)Anthropic67.3
73o3 Pro (batch)OpenAI67.3
74Claude Sonnet 5 (batch)Anthropic67.1
75GPT-5.2OpenAI66.6
76GPT-5.6 Luna ProOpenAI66.1
77GPT-5.6 Luna Pro (batch)OpenAI66.1
78GPT-5.6 LunaOpenAI66.1
79GPT-5.6 Luna (batch)OpenAI66.1
80Nemotron 3 Nano Omni (free)NVIDIA66
81OpenAI GPT Mini Latest~openai66
82GPT-5.4 MiniOpenAI66
83GPT-5.2-CodexOpenAI65.9
84GPT-5 ImageOpenAI65.7
85GPT-5.4 Mini (batch)OpenAI65.5
86Claude Sonnet 4.5Anthropic65.4
87GPT-5.4 NanoOpenAI65.2
88GPT-5.4 Nano (batch)OpenAI65.1
89Sakana Namazusakana65
90Anthropic Claude Haiku Latest~anthropic64.9
91GPT-5.2 (batch)OpenAI64.9
92GPT-5.1OpenAI64.9
93Gemini 2.5 ProGoogle64.7
94Claude Opus 4.5 (batch)Anthropic64.6
95MiMo-V2.5Xiaomi64.1
96Inklingthinkingmachines63.9
97Muse Spark 1.2meta63.8
98Qwen3.8 MaxAlibaba63.8
99Muse Spark 1.1meta63.8
100GPT-5.1-Codex-MaxOpenAI63.8
101GPT-5 Image MiniOpenAI63.8
102GPT-5.1 (batch)OpenAI63.7
103MiniMax M3MiniMax63.6
104Gemini 2.5 Pro Preview 05-06Google63.6
105Nano Banana 2 (Gemini 3.1 Flash Image)Google63.5
106Claude Sonnet 4.5 (batch)Anthropic63.5
107Gemini 2.5 Pro (batch)Google63.4
108Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google63.3
109GPT-5.1-CodexOpenAI63.2
110Gemini 2.5 Flash LiteGoogle63.2
111Gemini 2.5 Flash Lite (batch)Google63.2
112Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google62.9
113Gemini 2.5 FlashGoogle62.8
114Nano Banana Pro (Gemini 3 Pro Image Preview)Google62.7
115Gemini 2.5 Pro Preview 06-05Google62.7
116Gemini 2.5 Flash (batch)Google62.5
117Inkling Smallthinkingmachines62.4
118Qwen3.5 Plus 2026-04-20Alibaba62.3
119Qwen3.6 27BAlibaba62.3
120Qwen3.6 PlusAlibaba62.3
121Qwen3.5 397B A17BAlibaba62.3
122Lyria 3 Pro PreviewGoogle62.2
123Lyria 3 Clip PreviewGoogle62.2
124Qwen3.5 Plus 2026-02-15Alibaba62.2
125GPT-5OpenAI62.2
126Qwen3.6 FlashAlibaba62.1
127Seed 2.1 TurboByteDance62
128Qwen3.5-FlashAlibaba61.9
129Qwen3.7 FlashAlibaba61.8
130Seed-2.0-CodeByteDance61.7
131Step 3.7 FlashStepFun61.7
132Qwen3.6 35B A3BAlibaba61.7
133GLM 5V TurboZhipu AI61.7
134Qwen3.5-35B-A3BAlibaba61.7
135Claude Haiku 4.5Anthropic61.6
136Gemma 4 26B A4B Google61.5
137Gemma 4 31BGoogle61.5
138Qwen3.5-9BAlibaba61.5
139Seed-2.0-LiteByteDance61.4
140Qwen3.5-122B-A10BAlibaba61.2
141GPT-5.1-Codex-MiniOpenAI61.2
142Nemotron Nano 12B 2 VL (free)NVIDIA61.1
143Qwen3.7 PlusAlibaba61
144Seed-2.0-MiniByteDance61
145Claude Haiku 4.5 (batch)Anthropic61
146GPT-5 (batch)OpenAI61
147GPT-5.2 ChatOpenAI60.9
148Qwen3.5-27BAlibaba60.8
149Grok 4.20xAI60.7
150Kimi K2.7 CodeMoonshot AI60.6
151GPT-5 Codex (batch)OpenAI60.6
152Kimi K2.6Moonshot AI60.4
153Grok 4.6xAI60.2
154Grok 4.5xAI60.2
155Grok Latest~x-ai60.2
156GPT-5 MiniOpenAI60.2
157Grok 4.3xAI60
158Kimi K2.5Moonshot AI60
159GPT-5 Mini (batch)OpenAI60
160o1 (batch)OpenAI60
161GPT-5 NanoOpenAI59.8
162GPT-5 Nano (batch)OpenAI59.8
163Kimi K3Moonshot AI59.6
164Seed 1.6ByteDance59
165Seed 1.6 FlashByteDance58.6
166Grok Build 0.1xAI58.5
167Nova 2 LiteAmazon58.5
168Claude Sonnet 4Anthropic58.3
169Sonar Pro SearchPerplexity57.8
170o3OpenAI57.8
171GLM 4.6VZhipu AI57.6
172Grok 4.20 Multi-AgentxAI57.1
173Nemotron 3.5 Content Safety (free)NVIDIA56.9
174o4 Mini HighOpenAI56.9
175o4 MiniOpenAI56.9
176o3 (batch)OpenAI56.8
177Mistral Medium 3.5Mistral AI56.3
178o4 Mini High (batch)OpenAI56.3
179o4 Mini (batch)OpenAI56.3
180MiniMax M3 (batch)MiniMax55.3
181Qwen3 VL 30B A3B ThinkingAlibaba55.3
182GPT-4.1OpenAI55
183Qwen3 VL 8B ThinkingAlibaba54.7
184Qwen3 VL 235B A22B ThinkingAlibaba54.6
185GPT-4.1 (batch)OpenAI54
186Nano Banana (Gemini 2.5 Flash Image)Google53.9
187Perceptron Mk1perceptron53.6
188GPT-4.1 MiniOpenAI53.4
189Kimi K2.7 Code (batch)Moonshot AI53.3
190GPT-4.1 Mini (batch)OpenAI53.2
191GPT-4.1 NanoOpenAI53.1
192GPT-4.1 Nano (batch)OpenAI53.1
193Reka Edgerekaai53
194Mistral Small 4Mistral AI52.9
195Nova Premier 1.0Amazon52.7
196Muse Glimmer 30Bmeta52.4
197Inkling (batch)thinkingmachines52.1
198GLM 4.5VZhipu AI51.7
199Qwen3 VL 8B InstructAlibaba51.5
200Qwen3 VL 235B A22B InstructAlibaba51.1
201Qwen3 VL 32B InstructAlibaba51
202Qwen3 VL 30B A3B InstructAlibaba50.8
203Mistral Large 3 2512Mistral AI49.2
204Mistral Small 3.2 24BMistral AI47.7
205Llama 4 ScoutMeta47.4
206Ministral 3 14B 2512Mistral AI47.2
207Ministral 3 8B 2512Mistral AI47.2
208Gemma 3 27BGoogle46.6
209Ministral 3 3B 2512Mistral AI46.5
210Mistral Medium 3.1Mistral AI45.6
211GPT-4o (2024-11-20)OpenAI45.3
212GPT-4o (batch)OpenAI44.9
213GPT-4 TurboOpenAI44.9
214GPT-4 Turbo (batch)OpenAI44.8
215Gemma 3 12BGoogle44.3
216Llama Guard 4 12BMeta44.2
217Sonar ProPerplexity44.2
218ERNIE 4.5 VL 424B A47B Baidu43.9
219GPT-4o-mini (batch)OpenAI43.8
220Mistral Medium 3Mistral AI42.9
221GPT-4o (2024-05-13)OpenAI42.9
222UI-TARS 7B ByteDance42.8
223GPT-4o (2024-08-06)OpenAI42.6
224GPT-4oOpenAI42.6
225Llama 4 MaverickMeta42.2
226Sonar Reasoning ProPerplexity41.1
227MiniMax-01MiniMax40.9
228Gemma 3 4BGoogle40.7
229GPT-4o-miniOpenAI40.3
230GPT-4o-mini (2024-07-18)OpenAI40.3
231Nova Pro 1.0Amazon38.9
232Mistral Small 3.1 24BMistral AI38.8
233Claude 3 HaikuAnthropic38.2
234Nova Lite 1.0Amazon38.1
235Qwen2.5 VL 72B InstructAlibaba34.8
236SonarPerplexity34.7

What Else Can Vision Models Do?

Vision-capable models often support additional capabilities. Here is how the 236 vision models break down by other features.

Function Calling

89%

210 of 236 vision models

JSON Mode

92%

218 of 236 vision models

Streaming

100%

236 of 236 vision models

Reasoning

79%

187 of 236 vision models

Web Search

63%

149 of 236 vision models

Image Output

4%

9 of 236 vision models

Top Providers of Vision Models

Which providers offer the most AI models with image understanding capabilities.

OpenAI

77

vision models

Google

37

vision models

Anthropic

28

vision models

Alibaba

23

vision models

Mistral AI

10

vision models

ByteDance

7

vision models

xAI

6

vision models

Moonshot AI

5

vision models

Top Vision Models - Full Capability Breakdown

Detailed capability matrix for the top 15 vision models by composite score.

Supported
Not supported

What Are Vision / Multimodal AI Models?

Vision-capable AI models (also called multimodal models) can process both text and images as input. Instead of being limited to text-only prompts, these models can understand photos, screenshots, charts, diagrams, handwritten notes, and other visual content.

How Vision Input Works

You send an image (as a URL or base64-encoded data) alongside your text prompt via the API. The model processes both together, allowing it to answer questions about the image, describe its contents, extract text (OCR), analyze charts, or reason about visual information.

Common Use Cases for Vision Models

Document analysis and OCR, screenshot-to-code generation, chart and graph interpretation, medical image analysis, accessibility (image descriptions), visual QA for customer support, UI/UX review, diagram-to-code conversion, and real-time visual understanding in agentic workflows.

Vision vs. Image Output

Vision (image input) and image output are distinct capabilities. Vision means the model can see and understand images you send to it. Image output means the model can generate new images. Some models support both, but many vision models are text-output only - they analyze images but respond with text.

Pricing Considerations

Image tokens typically cost more than text tokens. A single high-resolution image can consume thousands of tokens. Most providers charge for image input based on the image dimensions and detail level. Check each model's specific pricing for image token costs.

Explore More Model Comparisons

Dive deeper into model capabilities, rankings, and head-to-head comparisons to find the right vision model for your use case.

Frequently Asked Questions

Vision AI models (multimodal LLMs) can understand and analyze images alongside text. They accept image inputs and can describe photos, extract text from screenshots, analyze charts, and answer questions about visual content.

The top vision models include GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. Rankings change frequently as models improve — check our leaderboard for the latest scores.

Not all of them. Vision capabilities refer to understanding images (input). Some models like DALL-E and Stable Diffusion specialize in generating images (output). A few models like GPT-4o can both understand and generate images.

AI Models with Vision - Best Multimodal AI (2026) | LM Market Cap