Eight leading models, one current ranking, and the practical details needed to choose among them. Claude Fable 5 leads this edition at 83.0, but only 5.0 points separate first from eighth. That narrow spread makes workload fit, price, access and deployment control decisive.

Quick answer: Claude Fable 5 is the benchmark leader. Kimi K3 is the highest-ranked open-weights option. GPT 5.6 Sol is the highest-ranked OpenAI model, while Claude Opus 5 is the lower-priced Anthropic alternative at a published $5 / $25 token price.

Edition snapshot: August 28, 2026

Decision

Current leader

Why it stands out

Highest Overall score

Claude Fable 5

#1 with 83.0

Highest-ranked open weights

Kimi K3

#5 with 79.2

Highest-ranked OpenAI model

GPT 5.6 Sol

#2 with 81.0

Lowest numeric proprietary price shown

Grok 4.6

$2 input / $6 output

Largest specified input window

GPT 5.6 Sol and GPT 5.5

1.05 million tokens

Complete AI model comparison

Rank

Model

Score

Best fit

Context or scale

API price

1

Claude Fable 5

83.0

Frontier reasoning, long-horizon coding and research

Long-running, million-token-class work

$10 input / $50 output

2

GPT 5.6 Sol

81.0

Reasoning, tool-using agents, coding and large repositories

1.05M input / 128K output

$4 input / $20 output, promotional

3

GPT 5.5

80.2

Professional work, coding, analysis and adjustable reasoning

1.05M input / 128K output

$5 input / $30 output

4

Claude Opus 5

80.1

Coding, computer use and analytical work

Long-running agentic work

$5 input / $25 output

5

Kimi K3

79.2

Open-weight frontier work, coding, research and self-hosting

1M tokens with native vision

$3 input / $15 output

6

Gemini 3.7 Flash

78.8

Fast multimodal agents and high-volume applications

1,048,576 input / 65,536 output

Usage-based; verify current Google rates

7

Qwen 3.8

78.5

Open-weight coding, multimodal agents and deployment control

2.4T parameters, 95B active

API and self-hosting rates vary

8

Grok 4.6

78.0

Cost-sensitive agents, coding and partner workflows

Agentic reasoning and knowledge work

From $2 input / $6 output

Token prices are USD per million tokens, input / output, based on the current model guides. Caching, batch, long-context, regional and tool charges can change the effective cost.

Overall score spread

Model

Overall

Position on a 75 to 85 display range

Claude Fable 5

83.0

█████████

GPT 5.6 Sol

81.0

███████

GPT 5.5

80.2

██████

Claude Opus 5

80.1

██████

Kimi K3

79.2

█████

Gemini 3.7 Flash

78.8

█████

Qwen 3.8

78.5

████

Grok 4.6

78.0

████

The bars use a restricted 75 to 85 display range to make small differences visible. They are not percentages and should always be read with the numeric score.

Segmented winners

A single overall rank cannot answer every buying question. These narrower calls use only the current eight-model dataset and the verified facts summarized in each model guide.

Category

Winner

Reason

Best overall

Claude Fable 5

Highest current Overall score at 83.0

Best open weights

Kimi K3

Highest-ranked open-weights model at #5

Best OpenAI model

GPT 5.6 Sol

Highest-ranked OpenAI entry at #2

Lower-priced Anthropic alternative

Claude Opus 5

80.1 Overall at a published $5 / $25 token price

Lowest illustrative blended token price

Grok 4.6

$3.33 on the documented 2:1 calculation below

Best fit for large multimodal inputs

Gemini 3.7 Flash

1,048,576-token input capacity plus Google multimodal tooling

Best fit for open-weight deployment control

Kimi K3

Top open-weights rank plus downloadable weights

Capability and API cost

The table keeps capability and token price side by side without turning price into a second capability score. Input and output prices are USD per million tokens. The illustrative blend assumes two input tokens for every output token: (2 × input price + output price) ÷ 3.

Model

Overall

Input

Output

Illustrative 2:1 blend

Pricing note

Claude Fable 5

83.0

$10

$50

$23.33

Standard published rate

GPT 5.6 Sol

81.0

$4

$20

$9.33

Documented promotional rate

GPT 5.5

80.2

$5

$30

$13.33

Standard published rate

Claude Opus 5

80.1

$5

$25

$11.67

Standard published rate

Kimi K3

79.2

$3

$15

$7.00

Cache-miss input rate

Gemini 3.7 Flash

78.8

Not comparable

Not comparable

Not calculated

Verify current provider pricing

Qwen 3.8

78.5

Not comparable

Not comparable

Not calculated

Verify current provider pricing

Grok 4.6

78.0

$2

$6

$3.33

Base variant rate

This is a token-price comparison, not cost per completed task. Reasoning length, caching, retries, tool calls, long-context multipliers, hosting and regional charges can reverse the apparent value order. Gemini 3.7 Flash and Qwen 3.8 are excluded from the blend because the current guides do not contain directly comparable numeric rates.

Build your shortlist by need

If you prioritize

Start with

Also compare

Highest current Overall score

Claude Fable 5

GPT 5.6 Sol

Open weights and deployment control

Kimi K3

Qwen 3.8

OpenAI tools and APIs

GPT 5.6 Sol

GPT 5.5

Anthropic capability at a lower token price

Claude Opus 5

Claude Fable 5

Large multimodal inputs and Google tooling

Gemini 3.7 Flash

Kimi K3

Low published proprietary base token price

Grok 4.6

GPT 5.6 Sol

What the ranking actually says

The Overall score is The AI Leaderboard's current comparative score. It averages tasks within seven evaluation categories, then gives those category averages equal weight. This limits the effect of one unusually strong category.

The top four models sit within 2.9 points, and the entire top eight fits inside 5.0 points. Rank therefore identifies a strong shortlist, not a universal winner. A lower-ranked model can be the better production choice when it offers cheaper inference, open weights, preferred tools or a deployment model your team can operate reliably.

The top AI models, explained

#1 Claude Fable 5

Anthropic | Proprietary | Overall score: 83.0

The current overall leader. It is the first model to shortlist when frontier capability matters most, provided its higher price and tighter safeguards fit the workload.

  • Best fit: Frontier reasoning, long-horizon coding and research.
  • Access: Claude API and first-party Claude products.
  • Context or scale: Long-running, million-token-class work.
  • Current pricing reference: $10 input / $50 output.

Read the complete Claude Fable 5 guide or check the official model source.

#2 GPT 5.6 Sol

OpenAI | Proprietary | Overall score: 81.0

The highest-ranked OpenAI model in this edition. It combines a near-leading score with a lower documented token price than GPT 5.5 during its promotional period.

  • Best fit: Reasoning, tool-using agents, coding and large repositories.
  • Access: OpenAI API alias and versioned snapshots.
  • Context or scale: 1.05M input / 128K output.
  • Current pricing reference: $4 input / $20 output, promotional.

Read the complete GPT 5.6 Sol guide or check the official model source.

#3 GPT 5.5

OpenAI | Proprietary | Overall score: 80.2

A high-ranking OpenAI model that remains close to the leaders. Compare it directly with GPT 5.6 Sol before choosing a default because their score and pricing order do not match.

  • Best fit: Professional work, coding, analysis and adjustable reasoning.
  • Access: OpenAI API aliases and snapshots.
  • Context or scale: 1.05M input / 128K output.
  • Current pricing reference: $5 input / $30 output.

Read the complete GPT 5.5 guide or check the official model source.

#4 Claude Opus 5

Anthropic | Proprietary | Overall score: 80.1

A leading Anthropic model at half Fable 5's published token price. It is the clearest Anthropic alternative when cost matters but demanding agentic work remains central.

  • Best fit: Coding, computer use and analytical work.
  • Access: Claude API and first-party Claude products.
  • Context or scale: Long-running agentic work.
  • Current pricing reference: $5 input / $25 output.

Read the complete Claude Opus 5 guide or check the official model source.

#5 Kimi K3

Moonshot AI | Open weights | Overall score: 79.2

The highest-ranked open-weights model in the current top eight. It offers deployment control, but its exceptionally large architecture makes efficient self-hosting an infrastructure decision, not a simple download.

  • Best fit: Open-weight frontier work, coding, research and self-hosting.
  • Access: Kimi products, API and downloadable weights.
  • Context or scale: 1M tokens with native vision.
  • Current pricing reference: $3 input / $15 output.

Read the complete Kimi K3 guide or check the official model source.

#6 Gemini 3.7 Flash

Google | Proprietary | Overall score: 78.8

The highest-ranked Google model in this edition. Its fit is strongest when large mixed-media inputs, search grounding or Google ecosystem integration matter alongside score.

  • Best fit: Fast multimodal agents and high-volume applications.
  • Access: Gemini API and Google AI Studio.
  • Context or scale: 1,048,576 input / 65,536 output.
  • Current pricing reference: Usage-based; verify current Google rates.

Read the complete Gemini 3.7 Flash guide or check the official model source.

#7 Qwen 3.8

Alibaba | Open weights | Overall score: 78.5

An open-weights Alibaba model for teams that value deployment control. Reproduction requires the exact Max checkpoint because the public family label intentionally removes configuration qualifiers.

  • Best fit: Open-weight coding, multimodal agents and deployment control.
  • Access: QwenCloud and announced open weights.
  • Context or scale: 2.4T parameters, 95B active.
  • Current pricing reference: API and self-hosting rates vary.

Read the complete Qwen 3.8 guide or check the official model source.

#8 Grok 4.6

xAI | Proprietary | Overall score: 78.0

xAI's current entry in the top eight and the lowest published base token price among the proprietary models with numeric prices shown here. Configuration and partner differences still need direct testing.

  • Best fit: Cost-sensitive agents, coding and partner workflows.
  • Access: xAI API, Grok Build and selected partners.
  • Context or scale: Agentic reasoning and knowledge work.
  • Current pricing reference: From $2 input / $6 output.

Read the complete Grok 4.6 guide or check the official model source.

Which AI model should you choose?

Start with the job, not the brand. Write down the tasks the model must complete, the information it can access, the actions it may take and the evidence a reviewer needs before accepting its output.

  • Choose two or three candidates from the table rather than testing every model.
  • Run the same representative tasks with the same prompts, tools and context.
  • Score task completion, factual support, instruction following, latency and total cost.
  • Test failure handling, not only ideal examples.
  • Keep a human approval step for consequential decisions and external actions.
  • Record the exact model version and settings so results remain reproducible.

Proprietary or open weights?

Decision factor

Proprietary models

Open-weights models

Setup

Usually fastest through a hosted API

Requires hosting or a specialist provider

Control

Provider controls the service and updates

More control over deployment and versioning

Operations

Provider manages infrastructure

Your team owns capacity, security and monitoring

Customization

Depends on provider features

Often supports deeper deployment customization

Best fit

Teams prioritizing speed and managed access

Teams prioritizing control, portability or self-hosting

Methodology and update policy

This page reflects the published leaderboard and its current verified scores. It does not create a separate composite or invent scores for models outside the ranked dataset. Read how we rank for the site's methodology and evidence policy.

Page element

Primary evidence

Important limitation

Overall score and rank

Current validated AI Leaderboard dataset

A benchmark score is not a production success rate

Model identity and provider

Canonical family mapping plus official provider source

Public labels remove effort and tier qualifiers

Access, context and pricing

Official model documentation summarized in each guide

Prices and availability can change between refreshes

Best-fit guidance

Editorial interpretation of documented capabilities and deployment model

Your own tasks may produce a different order

The directory is refreshed when the leaderboard is updated. Model names use canonical family labels, while configuration details such as reasoning or effort settings are retained during source validation. Reproducing a result requires the exact dated release, model variant and evaluation harness.

What this ranking does not measure

The published Overall score does not directly measure your private data, regional availability, integration quality, latency target, security controls or total production cost. Those factors belong in a workload-specific evaluation before procurement or migration.

Current edition notes

This page represents the leaderboard refreshed on August 28, 2026. It records the current order without claiming historical movement where no verified prior snapshot is shown. Future editions can add movement indicators once consecutive validated snapshots are retained.

Frequently asked questions

What is the best AI model right now?

Claude Fable 5 is number one on the current AI Leaderboard with an Overall score of 83.0. That makes it the present benchmark leader in this dataset, not the automatic best choice for every task.

What is the best open-weights AI model?

Kimi K3 is the highest-ranked open-weights model in the current top eight, at number 5 with an Overall score of 79.2.

Are the Overall scores percentages?

No. They are one-decimal comparative benchmark scores. A score of 83.0 does not mean that a model is correct 83 percent of the time.

Why are only eight models included?

The live directory currently covers the eight verified records maintained in The AI Leaderboard dataset. We expand coverage only when the model identity, source data and public guide can be validated.

How often is this page updated?

It follows the leaderboard refresh cycle and is updated as verified ranking data changes. The date near the top shows the edition represented on the page.