Eight leading models, one current ranking, and the practical details needed to choose among them. Claude Fable 5 leads this edition at 83.0, but only 5.0 points separate first from eighth. That narrow spread makes workload fit, price, access and deployment control decisive.
Quick answer: Claude Fable 5 is the benchmark leader. Kimi K3 is the highest-ranked open-weights option. GPT 5.6 Sol is the highest-ranked OpenAI model, while Claude Opus 5 is the lower-priced Anthropic alternative at a published $5 / $25 token price.
Edition snapshot: August 28, 2026
Decision | Current leader | Why it stands out |
|---|---|---|
Highest Overall score | Claude Fable 5 | #1 with 83.0 |
Highest-ranked open weights | Kimi K3 | #5 with 79.2 |
Highest-ranked OpenAI model | GPT 5.6 Sol | #2 with 81.0 |
Lowest numeric proprietary price shown | Grok 4.6 | $2 input / $6 output |
Largest specified input window | GPT 5.6 Sol and GPT 5.5 | 1.05 million tokens |
Complete AI model comparison
Rank | Model | Score | Best fit | Context or scale | API price |
|---|---|---|---|---|---|
1 | Claude Fable 5 | 83.0 | Frontier reasoning, long-horizon coding and research | Long-running, million-token-class work | $10 input / $50 output |
2 | GPT 5.6 Sol | 81.0 | Reasoning, tool-using agents, coding and large repositories | 1.05M input / 128K output | $4 input / $20 output, promotional |
3 | GPT 5.5 | 80.2 | Professional work, coding, analysis and adjustable reasoning | 1.05M input / 128K output | $5 input / $30 output |
4 | Claude Opus 5 | 80.1 | Coding, computer use and analytical work | Long-running agentic work | $5 input / $25 output |
5 | Kimi K3 | 79.2 | Open-weight frontier work, coding, research and self-hosting | 1M tokens with native vision | $3 input / $15 output |
6 | Gemini 3.7 Flash | 78.8 | Fast multimodal agents and high-volume applications | 1,048,576 input / 65,536 output | Usage-based; verify current Google rates |
7 | Qwen 3.8 | 78.5 | Open-weight coding, multimodal agents and deployment control | 2.4T parameters, 95B active | API and self-hosting rates vary |
8 | Grok 4.6 | 78.0 | Cost-sensitive agents, coding and partner workflows | Agentic reasoning and knowledge work | From $2 input / $6 output |
Token prices are USD per million tokens, input / output, based on the current model guides. Caching, batch, long-context, regional and tool charges can change the effective cost.
Overall score spread
Model | Overall | Position on a 75 to 85 display range |
|---|---|---|
Claude Fable 5 | 83.0 | █████████ |
GPT 5.6 Sol | 81.0 | ███████ |
GPT 5.5 | 80.2 | ██████ |
Claude Opus 5 | 80.1 | ██████ |
Kimi K3 | 79.2 | █████ |
Gemini 3.7 Flash | 78.8 | █████ |
Qwen 3.8 | 78.5 | ████ |
Grok 4.6 | 78.0 | ████ |
The bars use a restricted 75 to 85 display range to make small differences visible. They are not percentages and should always be read with the numeric score.
Segmented winners
A single overall rank cannot answer every buying question. These narrower calls use only the current eight-model dataset and the verified facts summarized in each model guide.
Category | Winner | Reason |
|---|---|---|
Best overall | Claude Fable 5 | Highest current Overall score at 83.0 |
Best open weights | Kimi K3 | Highest-ranked open-weights model at #5 |
Best OpenAI model | GPT 5.6 Sol | Highest-ranked OpenAI entry at #2 |
Lower-priced Anthropic alternative | Claude Opus 5 | 80.1 Overall at a published $5 / $25 token price |
Lowest illustrative blended token price | Grok 4.6 | $3.33 on the documented 2:1 calculation below |
Best fit for large multimodal inputs | Gemini 3.7 Flash | 1,048,576-token input capacity plus Google multimodal tooling |
Best fit for open-weight deployment control | Kimi K3 | Top open-weights rank plus downloadable weights |
Capability and API cost
The table keeps capability and token price side by side without turning price into a second capability score. Input and output prices are USD per million tokens. The illustrative blend assumes two input tokens for every output token: (2 × input price + output price) ÷ 3.
Model | Overall | Input | Output | Illustrative 2:1 blend | Pricing note |
|---|---|---|---|---|---|
Claude Fable 5 | 83.0 | $10 | $50 | $23.33 | Standard published rate |
GPT 5.6 Sol | 81.0 | $4 | $20 | $9.33 | Documented promotional rate |
GPT 5.5 | 80.2 | $5 | $30 | $13.33 | Standard published rate |
Claude Opus 5 | 80.1 | $5 | $25 | $11.67 | Standard published rate |
Kimi K3 | 79.2 | $3 | $15 | $7.00 | Cache-miss input rate |
Gemini 3.7 Flash | 78.8 | Not comparable | Not comparable | Not calculated | Verify current provider pricing |
Qwen 3.8 | 78.5 | Not comparable | Not comparable | Not calculated | Verify current provider pricing |
Grok 4.6 | 78.0 | $2 | $6 | $3.33 | Base variant rate |
This is a token-price comparison, not cost per completed task. Reasoning length, caching, retries, tool calls, long-context multipliers, hosting and regional charges can reverse the apparent value order. Gemini 3.7 Flash and Qwen 3.8 are excluded from the blend because the current guides do not contain directly comparable numeric rates.
Build your shortlist by need
If you prioritize | Start with | Also compare |
|---|---|---|
Highest current Overall score | Claude Fable 5 | GPT 5.6 Sol |
Open weights and deployment control | Kimi K3 | Qwen 3.8 |
OpenAI tools and APIs | GPT 5.6 Sol | GPT 5.5 |
Anthropic capability at a lower token price | Claude Opus 5 | Claude Fable 5 |
Large multimodal inputs and Google tooling | Gemini 3.7 Flash | Kimi K3 |
Low published proprietary base token price | Grok 4.6 | GPT 5.6 Sol |
What the ranking actually says
The Overall score is The AI Leaderboard's current comparative score. It averages tasks within seven evaluation categories, then gives those category averages equal weight. This limits the effect of one unusually strong category.
The top four models sit within 2.9 points, and the entire top eight fits inside 5.0 points. Rank therefore identifies a strong shortlist, not a universal winner. A lower-ranked model can be the better production choice when it offers cheaper inference, open weights, preferred tools or a deployment model your team can operate reliably.
The top AI models, explained
#1 Claude Fable 5
Anthropic | Proprietary | Overall score: 83.0
The current overall leader. It is the first model to shortlist when frontier capability matters most, provided its higher price and tighter safeguards fit the workload.
- Best fit: Frontier reasoning, long-horizon coding and research.
- Access: Claude API and first-party Claude products.
- Context or scale: Long-running, million-token-class work.
- Current pricing reference: $10 input / $50 output.
Read the complete Claude Fable 5 guide or check the official model source.
#2 GPT 5.6 Sol
OpenAI | Proprietary | Overall score: 81.0
The highest-ranked OpenAI model in this edition. It combines a near-leading score with a lower documented token price than GPT 5.5 during its promotional period.
- Best fit: Reasoning, tool-using agents, coding and large repositories.
- Access: OpenAI API alias and versioned snapshots.
- Context or scale: 1.05M input / 128K output.
- Current pricing reference: $4 input / $20 output, promotional.
Read the complete GPT 5.6 Sol guide or check the official model source.
#3 GPT 5.5
OpenAI | Proprietary | Overall score: 80.2
A high-ranking OpenAI model that remains close to the leaders. Compare it directly with GPT 5.6 Sol before choosing a default because their score and pricing order do not match.
- Best fit: Professional work, coding, analysis and adjustable reasoning.
- Access: OpenAI API aliases and snapshots.
- Context or scale: 1.05M input / 128K output.
- Current pricing reference: $5 input / $30 output.
Read the complete GPT 5.5 guide or check the official model source.
#4 Claude Opus 5
Anthropic | Proprietary | Overall score: 80.1
A leading Anthropic model at half Fable 5's published token price. It is the clearest Anthropic alternative when cost matters but demanding agentic work remains central.
- Best fit: Coding, computer use and analytical work.
- Access: Claude API and first-party Claude products.
- Context or scale: Long-running agentic work.
- Current pricing reference: $5 input / $25 output.
Read the complete Claude Opus 5 guide or check the official model source.
#5 Kimi K3
Moonshot AI | Open weights | Overall score: 79.2
The highest-ranked open-weights model in the current top eight. It offers deployment control, but its exceptionally large architecture makes efficient self-hosting an infrastructure decision, not a simple download.
- Best fit: Open-weight frontier work, coding, research and self-hosting.
- Access: Kimi products, API and downloadable weights.
- Context or scale: 1M tokens with native vision.
- Current pricing reference: $3 input / $15 output.
Read the complete Kimi K3 guide or check the official model source.
#6 Gemini 3.7 Flash
Google | Proprietary | Overall score: 78.8
The highest-ranked Google model in this edition. Its fit is strongest when large mixed-media inputs, search grounding or Google ecosystem integration matter alongside score.
- Best fit: Fast multimodal agents and high-volume applications.
- Access: Gemini API and Google AI Studio.
- Context or scale: 1,048,576 input / 65,536 output.
- Current pricing reference: Usage-based; verify current Google rates.
Read the complete Gemini 3.7 Flash guide or check the official model source.
#7 Qwen 3.8
Alibaba | Open weights | Overall score: 78.5
An open-weights Alibaba model for teams that value deployment control. Reproduction requires the exact Max checkpoint because the public family label intentionally removes configuration qualifiers.
- Best fit: Open-weight coding, multimodal agents and deployment control.
- Access: QwenCloud and announced open weights.
- Context or scale: 2.4T parameters, 95B active.
- Current pricing reference: API and self-hosting rates vary.
Read the complete Qwen 3.8 guide or check the official model source.
#8 Grok 4.6
xAI | Proprietary | Overall score: 78.0
xAI's current entry in the top eight and the lowest published base token price among the proprietary models with numeric prices shown here. Configuration and partner differences still need direct testing.
- Best fit: Cost-sensitive agents, coding and partner workflows.
- Access: xAI API, Grok Build and selected partners.
- Context or scale: Agentic reasoning and knowledge work.
- Current pricing reference: From $2 input / $6 output.
Read the complete Grok 4.6 guide or check the official model source.
Which AI model should you choose?
Start with the job, not the brand. Write down the tasks the model must complete, the information it can access, the actions it may take and the evidence a reviewer needs before accepting its output.
- Choose two or three candidates from the table rather than testing every model.
- Run the same representative tasks with the same prompts, tools and context.
- Score task completion, factual support, instruction following, latency and total cost.
- Test failure handling, not only ideal examples.
- Keep a human approval step for consequential decisions and external actions.
- Record the exact model version and settings so results remain reproducible.
Proprietary or open weights?
Decision factor | Proprietary models | Open-weights models |
|---|---|---|
Setup | Usually fastest through a hosted API | Requires hosting or a specialist provider |
Control | Provider controls the service and updates | More control over deployment and versioning |
Operations | Provider manages infrastructure | Your team owns capacity, security and monitoring |
Customization | Depends on provider features | Often supports deeper deployment customization |
Best fit | Teams prioritizing speed and managed access | Teams prioritizing control, portability or self-hosting |
Methodology and update policy
This page reflects the published leaderboard and its current verified scores. It does not create a separate composite or invent scores for models outside the ranked dataset. Read how we rank for the site's methodology and evidence policy.
Page element | Primary evidence | Important limitation |
|---|---|---|
Overall score and rank | Current validated AI Leaderboard dataset | A benchmark score is not a production success rate |
Model identity and provider | Canonical family mapping plus official provider source | Public labels remove effort and tier qualifiers |
Access, context and pricing | Official model documentation summarized in each guide | Prices and availability can change between refreshes |
Best-fit guidance | Editorial interpretation of documented capabilities and deployment model | Your own tasks may produce a different order |
The directory is refreshed when the leaderboard is updated. Model names use canonical family labels, while configuration details such as reasoning or effort settings are retained during source validation. Reproducing a result requires the exact dated release, model variant and evaluation harness.
What this ranking does not measure
The published Overall score does not directly measure your private data, regional availability, integration quality, latency target, security controls or total production cost. Those factors belong in a workload-specific evaluation before procurement or migration.
Current edition notes
This page represents the leaderboard refreshed on August 28, 2026. It records the current order without claiming historical movement where no verified prior snapshot is shown. Future editions can add movement indicators once consecutive validated snapshots are retained.
Frequently asked questions
What is the best AI model right now?
Claude Fable 5 is number one on the current AI Leaderboard with an Overall score of 83.0. That makes it the present benchmark leader in this dataset, not the automatic best choice for every task.
What is the best open-weights AI model?
Kimi K3 is the highest-ranked open-weights model in the current top eight, at number 5 with an Overall score of 79.2.
Are the Overall scores percentages?
No. They are one-decimal comparative benchmark scores. A score of 83.0 does not mean that a model is correct 83 percent of the time.
Why are only eight models included?
The live directory currently covers the eight verified records maintained in The AI Leaderboard dataset. We expand coverage only when the model identity, source data and public guide can be validated.
How often is this page updated?
It follows the leaderboard refresh cycle and is updated as verified ranking data changes. The date near the top shows the edition represented on the page.