Skip to content

AI Models with Vision

284 multimodal AI models that support image input and visual understanding, ranked by composite score. These models can analyze images, screenshots, diagrams, and documents alongside text prompts.

Vision Models

284

Of All Models

61%

Providers

37

Avg Score

62

Vision Model Rankings

All 284 models with vision capability, sorted by composite score. Score is computed from capabilities, pricing, context window, recency, output capacity, and versatility.

#ModelScore
1GPT-5.5 ProOpenAI90.6
2GPT-5.4 ProOpenAI90
3GPT-5.5 Pro (batch)OpenAI88.1
4GPT-5.4 Pro (batch)OpenAI87.5
5GPT-5.2 ProOpenAI86.6
6GPT-5 ProOpenAI84.9
7GPT-5.2 Pro (batch)OpenAI82.6
8GPT Astra Latest~openai78.1
9GPT-6 AstraOpenAI78.1
10GPT-6 Astra ProOpenAI78.1
11Claude Fable 5.1Anthropic78
12Claude Fable Latest~anthropic78
13Claude Fable 5Anthropic78
14o3 ProOpenAI75.8
15GPT-5 Pro (batch)OpenAI74.9
16o1-proOpenAI74.9
17GPT-5.5OpenAI73.1
18Fugu Ultra v2sakana73
19GPT-6 Astra (batch)OpenAI71.8
20GPT-6 Astra Pro (batch)OpenAI71.8
21Claude Fable 5.1 (batch)Anthropic71.8
22Claude Opus 5Anthropic71.8
23Claude Fable 5 (batch)Anthropic71.8
24Claude Opus 4.8Anthropic71.8
25Claude Opus 4.7Anthropic71.8
26Claude Opus 4.1Anthropic71.7
27Gemini Pro Latest~google71.5
28Muse Spark 1.3meta71.4
29Muse Spark 1.2meta71.4
30Muse Spark 1.1meta71.4
31Fugu Ultrasakana71.4
32Gemini 3.5 FlashGoogle70.7
33Gemini 3.1 Pro Preview Custom ToolsGoogle70.6
34Space Bunny Alphastealth70.5
35Claude Opus 5.5Anthropic70.5
36Claude Opus Latest~anthropic70.5
37Muse Spark 1.3 Contributormeta70.4
38Muse Spark 1.2 Contributormeta70.4
39Gemini 3.1 Pro PreviewGoogle70.4
40Claude Opus 4.6Anthropic70.4
41Gemini 3.5 Flash (batch)Google69.6
42GPT-5.4 Image 2OpenAI69.6
43Gemini 3.8 FlashGoogle69.4
44Gemini 3.7 FlashGoogle69.4
45Gemini 3.6 FlashGoogle69.4
46Gemini Flash Latest~google69.4
47GPT-5.5 (batch)OpenAI69.3
48Gemini 3.5 Flash LiteGoogle69.1
49Nano Banana Pro (Gemini 3 Pro Image)Google69.1
50Gemini 3.8 Flash (batch)Google68.9
51Gemini 3.7 Flash (batch)Google68.9
52Gemini 3.6 Flash (batch)Google68.9
53Gemini 3.1 Pro Preview (batch)Google68.9
54Gemini 3.5 Flash Lite (batch)Google68.8
55Gemini 3.1 Flash LiteGoogle68.8
56Claude Opus 5 (batch)Anthropic68.7
57Claude Opus 4.8 (batch)Anthropic68.7
58Claude Opus 4.7 (batch)Anthropic68.7
59Grok 4.20xAI68.7
60GPT-5.4OpenAI68.7
61GPT Terra Latest~openai68.6
62GPT-5.6 Terra ProOpenAI68.6
63GPT-5.6 TerraOpenAI68.6
64Gemini 3.1 Flash Lite (batch)Google68.6
65Qwen3.8 27B (free)Alibaba68.5
66GPT Chat LatestOpenAI68.5
67Claude Sonnet 4.6Anthropic68.2
68GPT-6 Sol ProOpenAI68.1
69GPT-6 SolOpenAI68.1
70GPT Sol Latest~openai68.1
71GPT-5.6 Sol ProOpenAI68.1
72GPT-5.6 SolOpenAI68.1
73Gemini 3.1 Flash Lite PreviewGoogle68.1
74Claude Opus 5.5 (batch)Anthropic68
75Dots3-Note Preview (free)dots-studio68
76Claude Sonnet 5Anthropic68
77Claude Sonnet Latest~anthropic68
78Grok 4.3xAI67.5
79Grok 4.3 (batch)xAI67.4
80Grok 4.6xAI67.2
81Kimi K3Moonshot AI67.2
82Grok 4.5xAI67.2
83GPT-5.3-CodexOpenAI67.2
84Claude Opus 4.6 (batch)Anthropic67.2
85GPT-5.6 Terra Pro (batch)OpenAI67.1
86GPT-5.6 Terra (batch)OpenAI67.1
87Gemma 4 26B A4B (free)Google67.1
88Gemma 4 31B (free)Google67.1
89Fugu Maxsakana67
90Grok 4.7xAI66.9
91Grok Latest~x-ai66.9
92GPT-6 Sol Pro (batch)OpenAI66.8
93GPT-6 Sol (batch)OpenAI66.8
94GPT-5.6 Sol Pro (batch)OpenAI66.8
95GPT-5.6 Sol (batch)OpenAI66.8
96Claude Sonnet 5 (batch)Anthropic66.8
97GPT-5.4 (batch)OpenAI66.8
98Inkling Small (free)thinkingmachines66.4
99Inkling (free)thinkingmachines66.4
100Gemini 3 Flash PreviewGoogle66.4
101Claude Sonnet 4.6 (batch)Anthropic66.3
102Claude Opus 4.5Anthropic66.3
103Gemini 3 Flash Preview (batch)Google66.1
104Kimi Latest~moonshotai66
105o1OpenAI66
106MiMo-V2.6-Pro-UltraSpeedXiaomi65.9
107GPT-5.6 Luna ProOpenAI65.9
108GPT-5.6 LunaOpenAI65.9
109Claude Opus 4.1 (batch)Anthropic65.9
110GPT-6 Luna ProOpenAI65.7
111GPT-6 Luna Pro (batch)OpenAI65.7
112GPT-6 LunaOpenAI65.7
113GPT-6 Luna (batch)OpenAI65.7
114GPT Luna Latest~openai65.7
115GPT-5.6 Luna Pro (batch)OpenAI65.7
116GPT-5.6 Luna (batch)OpenAI65.7
117Nemotron 3 Nano Omni (free)NVIDIA65.7
118GPT Mini Latest~openai65.7
119Ember-1fireworks65.5
120GPT-5.4 MiniOpenAI65.4
121Grok 4.20 Multi-AgentxAI65.1
122GPT-5.2OpenAI65.1
123Qwen3.8 Max PrimeAlibaba65
124Grok Build 0.1xAI65
125GPT-5.4 Mini (batch)OpenAI64.8
126Sakana Namazusakana64.7
127Claude Haiku Latest~anthropic64.6
128GPT-5.4 NanoOpenAI64.6
129GPT-5.4 Nano (batch)OpenAI64.4
130GPT-5.2-CodexOpenAI64.4
131GPT-5 ImageOpenAI64.2
132MiMo-V2.6-ProXiaomi63.9
133Claude Sonnet 4.5Anthropic63.9
134MiMo-V2.6-FlashXiaomi63.8
135Qwen3.8 Omni FlashAlibaba63.8
136MiMo-V2.5Xiaomi63.8
137Qwen3.8 Max (0902)Alibaba63.5
138Kimi K3 (batch)Moonshot AI63.4
139GPT-5.2 (batch)OpenAI63.4
140GPT-5.1OpenAI63.4
141MiniMax M3MiniMax63.3
142Nano Banana 2 (Gemini 3.1 Flash Image)Google63.2
143Gemini 2.5 ProGoogle63.2
144Claude Opus 4.5 (batch)Anthropic63.1
145Mistral Medium 3.5Mistral AI62.8
146Qwen3.8 27BAlibaba62.7
147Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Google62.6
148GLM 5.3 FlashXZhipu AI62.3
149GLM Flash Latest~z-ai62.3
150GLM 5.3 FlashZhipu AI62.3
151GPT-5.1-Codex-MaxOpenAI62.3
152GPT-5 Image MiniOpenAI62.3
153Qwen3.8 FlashAlibaba62.1
154GLM 5.3 Flash (batch)Zhipu AI62.1
155Nano Banana 2 (Gemini 3.1 Flash Image Preview)Google62.1
156GPT-5.1 (batch)OpenAI62.1
157Qwen3.5 Plus 2026-04-20Alibaba62
158Qwen3.6 PlusAlibaba62
159Claude Sonnet 4.5 (batch)Anthropic62
160DeepSeek Flash Latest~deepseek61.9
161DeepSeek V4 Flash Vision ExpDeepSeek61.9
162Mistral Medium 3.5 (batch)Mistral AI61.9
163Lyria 3 Pro PreviewGoogle61.9
164Lyria 3 Clip PreviewGoogle61.9
165Gemini 2.5 Pro (batch)Google61.9
166Qwen3.6 FlashAlibaba61.8
167Qwen3.6 27BAlibaba61.8
168Gemini 2.5 Flash LiteGoogle61.8
169GPT-5.1-CodexOpenAI61.7
170Gemini 2.5 Flash Lite (batch)Google61.7
171Seed 2.1 TurboByteDance61.6
172Qwen3.7 FlashAlibaba61.5
173DeepSeek V4.1 FlashDeepSeek61.4
174Seed-2.0-CodeByteDance61.4
175Step 3.7 FlashStepFun61.3
176Qwen3.6 35B A3BAlibaba61.3
177GLM 5V TurboZhipu AI61.3
178Nano Banana Pro (Gemini 3 Pro Image Preview)Google61.3
179Gemini 2.5 FlashGoogle61.3
180Gemini 2.5 Pro Preview 06-05Google61.2
181Gemma 4 26B A4B Google61.1
182Gemini 2.5 Flash (batch)Google61
183Qwen3.5 Plus 2026-02-15Alibaba60.8
184Qwen3.5 397B A17BAlibaba60.8
185Qwen3.5-FlashAlibaba60.7
186GPT-5OpenAI60.7
187Qwen3.7 PlusAlibaba60.6
188Seed-2.0-LiteByteDance60.6
189DeepSeek V4.1 Flash (batch)DeepSeek60.5
190Inklingthinkingmachines60.4
191Kimi K2.6Moonshot AI60.3
192Kimi K2.7 CodeMoonshot AI60.2
193Claude Haiku 4.5Anthropic60.2
194Seed-2.0-MiniByteDance59.9
195Qwen3.5-122B-A10BAlibaba59.8
196Ling 3.0 Flash VLinclusionai59.7
197Qwen3.5-27BAlibaba59.7
198GPT-5.1-Codex-MiniOpenAI59.7
199Claude Haiku 4.5 (batch)Anthropic59.5
200GPT-5 (batch)OpenAI59.5
201GPT-5.2 ChatOpenAI59.4
202Inkling Smallthinkingmachines59.2
203Gemma 4 31BGoogle59.2
204Qwen3.5-9BAlibaba59.2
205Mistral Small 4Mistral AI59.1
206Mistral Small 4 (batch)Mistral AI59
207GPT-5 MiniOpenAI58.7
208Qwen3.5-35B-A3BAlibaba58.6
209Command A+Cohere58.5
210GPT-5 Mini (batch)OpenAI58.5
211GPT-5 NanoOpenAI58.3
212GPT-5 Nano (batch)OpenAI58.3
213Kimi K2.5Moonshot AI58.2
214Ternary Bonsai 2 27Bprism-ml58.1
215Seed 1.6ByteDance57.5
216Muse Glimmer 30Bmeta57.1
217Seed 1.6 FlashByteDance57.1
218Nova 2 LiteAmazon57
219Nemotron 3.5 Content Safety (free)NVIDIA56.6
220Sonar Pro SearchPerplexity56.4
221o3OpenAI56.3
222GLM 4.6VZhipu AI56.1
223o4 Mini HighOpenAI55.4
224o4 MiniOpenAI55.4
225o3 (batch)OpenAI55.3
226Claude Sonnet 4Anthropic55.1
227o4 Mini (batch)OpenAI54.8
228Mistral Large 3 2512Mistral AI54.5
229Mistral Large 3 2512 (batch)Mistral AI54.3
230Qwen3 VL 30B A3B ThinkingAlibaba53.8
231Paretounbiased53.7
232GPT-4.1OpenAI53.6
233Perceptron Mk1perceptron53.3
234Qwen3 VL 8B ThinkingAlibaba53.3
235Qwen3 VL 235B A22B ThinkingAlibaba53.2
236Ministral 3 14B 2512Mistral AI52.6
237GPT-4.1 (batch)OpenAI52.6
238Ministral 3 8B 2512Mistral AI52.5
239Ministral 3 8B 2512 (batch)Mistral AI52.5
240Nano Banana (Gemini 2.5 Flash Image)Google52.5
241Reka Edgerekaai52.4
242GPT-4.1 MiniOpenAI52
243GPT-4.1 Mini (batch)OpenAI51.8
244GPT-4.1 NanoOpenAI51.7
245GPT-4.1 Nano (batch)OpenAI51.6
246Ministral 3 3B 2512Mistral AI51.3
247Nova Premier 1.0Amazon51.3
248Nemotron 3.5 Content SafetyNVIDIA51.1
249Mistral Medium 3.1Mistral AI50.5
250Mistral Medium 3.1 (batch)Mistral AI50.2
251GLM 4.5VZhipu AI50.2
252Qwen3 VL 8B InstructAlibaba50
253Qwen3 VL 235B A22B InstructAlibaba49.8
254Qwen3 VL 32B InstructAlibaba49.5
255Qwen3 VL 30B A3B InstructAlibaba49.3
256Mistral Medium 3Mistral AI47.8
257Mistral Small 3.2 24BMistral AI46.3
258Sonar Reasoning ProPerplexity46
259Llama 4 ScoutMeta45.9
260Llama 4 MaverickMeta45.8
261GPT-4o (batch)OpenAI44.7
262GPT-4 TurboOpenAI44.7
263GPT-4 Turbo (batch)OpenAI44.5
264Gemma 3 27BGoogle44.3
265GPT-4o (2024-11-20)OpenAI43.9
266GPT-4o-mini (batch)OpenAI43.5
267Gemma 3 12BGoogle42.9
268Sonar ProPerplexity42.8
269GPT-4o (2024-05-13)OpenAI42.6
270ERNIE 4.5 VL 424B A47B Baidu42.5
271GPT-4o (2024-08-06)OpenAI42.3
272GPT-4oOpenAI42.3
273UI-TARS 7B ByteDance41.4
274Mistral Small 3.1 24BMistral AI40.8
275GPT-4o-miniOpenAI40
276GPT-4o-mini (2024-07-18)OpenAI40
277Qwen2.5 VL 72B InstructAlibaba39.7
278SonarPerplexity39.6
279Gemma 3 4BGoogle39.3
280MiniMax-01MiniMax39.3
281Claude 3 HaikuAnthropic38
282Nova Pro 1.0Amazon37.4
283Llama Guard 4 12BMeta37.2
284Nova Lite 1.0Amazon36.7

What Else Can Vision Models Do?

Vision-capable models often support additional capabilities. Here is how the 284 vision models break down by other features.

Function Calling

91%

259 of 284 vision models

JSON Mode

92%

262 of 284 vision models

Streaming

100%

284 of 284 vision models

Reasoning

81%

231 of 284 vision models

Web Search

60%

169 of 284 vision models

Image Output

3%

9 of 284 vision models

Top Providers of Vision Models

Which providers offer the most AI models with image understanding capabilities.

OpenAI

84

vision models

Google

40

vision models

Alibaba

28

vision models

Anthropic

28

vision models

Mistral AI

15

vision models

xAI

8

vision models

ByteDance

7

vision models

Zhipu AI

6

vision models

Top Vision Models - Full Capability Breakdown

Detailed capability matrix for the top 15 vision models by composite score.

Supported
Not supported

What Are Vision / Multimodal AI Models?

Vision-capable AI models (also called multimodal models) can process both text and images as input. Instead of being limited to text-only prompts, these models can understand photos, screenshots, charts, diagrams, handwritten notes, and other visual content.

How Vision Input Works

You send an image (as a URL or base64-encoded data) alongside your text prompt via the API. The model processes both together, allowing it to answer questions about the image, describe its contents, extract text (OCR), analyze charts, or reason about visual information.

Common Use Cases for Vision Models

Document analysis and OCR, screenshot-to-code generation, chart and graph interpretation, medical image analysis, accessibility (image descriptions), visual QA for customer support, UI/UX review, diagram-to-code conversion, and real-time visual understanding in agentic workflows.

Vision vs. Image Output

Vision (image input) and image output are distinct capabilities. Vision means the model can see and understand images you send to it. Image output means the model can generate new images. Some models support both, but many vision models are text-output only - they analyze images but respond with text.

Pricing Considerations

Image tokens typically cost more than text tokens. A single high-resolution image can consume thousands of tokens. Most providers charge for image input based on the image dimensions and detail level. Check each model's specific pricing for image token costs.

Explore More Model Comparisons

Dive deeper into model capabilities, rankings, and head-to-head comparisons to find the right vision model for your use case.

Frequently Asked Questions

Vision AI models (multimodal LLMs) can understand and analyze images alongside text. They accept image inputs and can describe photos, extract text from screenshots, analyze charts, and answer questions about visual content.

The top vision models include GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. Rankings change frequently as models improve — check our leaderboard for the latest scores.

Not all of them. Vision capabilities refer to understanding images (input). Some models like DALL-E and Stable Diffusion specialize in generating images (output). A few models like GPT-4o can both understand and generate images.

AI Models with Vision - Best Multimodal AI (2026) | LM Market Cap