LLM Leaderboard - AI Models Ranked by Benchmarks & Price
Top Performers by Task
Stay up to date with AI
Weekly rankings, new model releases, and pricing changes - straight to your inbox.
No spam. Unsubscribe anytime.
Based on our live composite scoring across benchmarks, pricing, speed, and capabilities, Claude Fable 5.1 currently leads the AI model leaderboard with a score of 96. Rankings are updated hourly using real-time data from 64+ providers.
We use a composite scoring system (0-100) that combines: benchmark performance (90%) from MMLU, GPQA, HumanEval, SWE-bench, and 15+ standardized evaluations, with capabilities and context window as tiebreakers (10%). Data is sourced from provider catalogs, benchmark feeds, and pricing sources, with page freshness shown from the underlying cache.
We track 468 AI models across 64 providers including OpenAI, Anthropic, Google, DeepSeek, Meta, and Mistral. This includes 440 coding models, 18 image generation models, and 10 video generation models.
Our coding category ranks 440 models by their performance on coding benchmarks like SWE-bench, HumanEval, and real-world coding tasks. The top coding models are recalculated as benchmark and pricing sources refresh. Visit our coding leaderboard for the current top 10.
Yes, we track 44 free AI models. These include models from Google, Meta, and other providers that offer free API access. Visit our free AI models page for the complete ranked list.