High confidence
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
72
Score Trend
72/100
14-day history
API Pricing
$0.58/M in
$1.44/M out
Context Window
524.3K
262.1K max output
#105range #89-#121Top 33%
#1#314
Signal Overview
Score Breakdown
| Signal | Strength | Weight | Impact |
|---|---|---|---|
| Benchmarksjust now | 71 | 30% | +21.2 |
| Recencyjust now | 100 | 15% | +15.0 |
| Pricingjust now | 99 | 15% | +14.8 |
| Capabilitiesjust now | 67 | 20% | +13.3 |
| Context Windowjust now | 91 | 10% | +9.1 |
| Output Capacityjust now | 90 | 10% | +9.0 |
Benchmark Performance
Benchmark Scores(0 benchmarks + Arena Elo)
LMSYS Arena Elo
1431
Percentile
88.5
Weight
30%
No task benchmark data available yet for this model.
Capabilities
Reasoning
Vision
Function Calling
JSON Mode
Streaming
Web Search
Image Output
Modalities
Input
text
image
audio
Output
text
Recent thinkingmachines releases
View this model against the provider’s recent shipping cadence.
Reviews
Community and practitioner feedback adds real-world signal on top of benchmarks and pricing.
Reviews
Be the first to review this model
Share your experience with Inkling Small and help the community make better decisions.
Frequently Asked Questions
Inkling Small by thinkingmachines excels in the Coding category, where it ranks #105 with a composite score of 72/100. Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of... It is particularly strong in areas highlighted by its top benchmark performance and adoption metrics, making it suitable for both individual developers and enterprise teams looking for a reliable coding solution.
Inkling Small is priced at $0.58 per million input tokens and $1.44 per million output tokens (USD). Contact the provider for volume discounts and enterprise pricing. Pricing is competitive within the coding category and reflects the model's quality-to-cost ratio.
In the Coding category, Inkling Small holds rank #105 out of 314 models tracked. Its quality rank is #105 and adoption rank is #105. You can use our comparison tool at /compare to see detailed side-by-side metrics with specific alternatives. Key differentiators include its composite scoring across benchmarks, community sentiment, and real-world adoption rates.
Inkling Small has been evaluated across 6 different signals. Its strongest areas include Capabilities (67/100), Benchmarks (71/100), Pricing (99/100). These scores are derived from industry-standard benchmarks, community ratings, and real-world performance metrics. The composite score of 72/100 reflects a weighted combination of all tracked signals.
Inkling Small is a paid model, though some providers may offer trial credits or limited free tiers for evaluation. Check thinkingmachines's website for current free tier availability and promotional offers.
Inkling Small supports a 524K token context window (524,288 tokens total). That translates to roughly 393,216 words in a single prompt. This is large enough to process entire codebases, research papers, or long conversation histories in one shot.
Inkling Small can generate up to 262K output tokens (262,144 tokens) per response. That is roughly 196,608 words. This is enough for generating complete code files, detailed reports, or long-form content in a single response.
Inkling Small supports image understanding (vision), function/tool calling, extended reasoning/chain-of-thought, streaming responses. Function calling lets you integrate it with external APIs and tools programmatically. Vision support means it can analyze images, screenshots, and diagrams alongside text. These capabilities determine which workflows and integrations the model can handle natively.
Yes, Inkling Small is an open-source model. You can download the weights, run it locally, fine-tune it for your use case, or deploy it on your own infrastructure. Many cloud providers also offer hosted versions if you prefer not to manage the infrastructure yourself. Self-hosting gives you full control over data privacy and eliminates per-token API costs.
Inkling Small was developed by thinkingmachines. It was released on July 30, 2026. You can access it through thinkingmachines's API or download the model weights directly. Check our provider page for all models from thinkingmachines and how they compare against each other.
Pick Inkling Small when you need a solid balance of cost and capability for everyday development tasks, content generation, and standard API integrations. If your task is straightforward text completion or classification, a cheaper model might give you 90% of the quality at a fraction of the price. Run a quick benchmark on your actual use case before committing.
You can access Inkling Small through thinkingmachines's API using standard HTTP requests or their official SDK. Most providers support OpenAI-compatible endpoints, so switching between models often requires changing just the model name in your API call. Streaming is supported for real-time token-by-token output. For production use, implement proper error handling, rate limiting, and cost monitoring.
Key Info
Providerthinkingmachines
CategoryCoding
Max Output262.1K tokens
LicenseOpen Source
Statusstable
HuggingFaceInkling-Small
Benchmark Scores(0 benchmarks + Arena Elo)
LMSYS Arena Elo
1431
Percentile
88.5
Weight
30%
No task benchmark data available yet for this model.
Data updated: Jul 31, 2026Benchmarks: Jul 31, 2026
Pricing Tools
Pricingper 1M tokens
Best value
81% cheaper than category average
Input
$0.58
-76% vs avg
Output
$1.44
-87% vs avg
Cost Estimator
Input: 70%Output: 30%
Est. monthly cost$8.38
Category average$49.87
You save $41.49/month vs category average
Access & Availability
Why This Rank
+Benchmarks
+Recency
+Pricing
+Capabilities