What is GPT-6 Sol?
GPT-6 Sol is OpenAI’s proprietary model for complex coding and agentic workflows. Released September 22, 2026 alongside GPT-6 Luna, it extended techniques from GPT-6 Astra into a lower-cost tier priced at $2 per million input tokens and $10 per million output tokens.
OpenAI released GPT-6.1 Sol one week later. GPT-6 Sol therefore matters mainly as an existing production model, migration reference and reproducibility target. New applications should compare it with GPT-6.1 Sol, which keeps the same base input and output rates while improving capability and cached-input pricing.
QUICK VERDICT: GPT-6 Sol brought GPT-6 coding and agentic capabilities to a much lower price tier than Astra. It remains available and capable, but GPT-6.1 Sol is now the newer model at the same base input and output rates with cheaper cached input, so most new Sol deployments should evaluate 6.1 first.
Field | Verified value |
|---|---|
Provider | OpenAI |
Release date | September 22, 2026 |
License | Proprietary |
API model ID | gpt-6-sol |
Context window | 1,050,000 tokens |
Maximum input | 922,000 tokens |
Maximum output | 128,000 tokens |
Modalities | Text and image input; text output |
Knowledge cutoff | April 20, 2026 |
Reasoning effort | None, low, medium, high, xhigh and max; medium is default |
Base API pricing | $2 input, $0.20 cached input and $10 output per million tokens |
Best fit | Complex coding and tool-using agent workflows |
Last verified | October 2, 2026 |
What changed from GPT-5.6 Sol
GPT-6 Sol lowered the cost of the Sol tier by half relative to GPT-5.6 promotional pricing while adding GPT-6 improvements in professional work, factuality, coding, computer use, communication style and alignment. It also introduced improved prompt caching and more explicit cache diagnostics.
Area | GPT-6 Sol change |
|---|---|
Input price | Reduced from $4 to $2 per million tokens |
Output price | Reduced from $20 to $10 per million tokens |
Factuality | OpenAI reports roughly half as many mistakes as GPT-5.6 Sol on its internal test |
Coding | Higher reported results on FrontierCode and DeepSWE |
Computer use | More cost-efficient performance than the previous Sol generation |
Caching | Higher hit rates, diagnostics, breakpoints and cache reuse across effort changes |
The practical question is not whether GPT-6 Sol is newer. It is whether the change improves accepted-task quality, speed and cost in the exact workflow. Keep prompts, tools, source material and scoring criteria fixed when comparing it with GPT-5.6 Sol.
Capabilities and best use cases
GPT-6 Sol supports the Responses API tool stack, including web search, file search, code execution, hosted shell, computer use and MCP connections. It can also accept text and images. The model is designed for sustained agent work where Astra’s price is unnecessary but Luna may not clear the quality threshold.
Capability | Practical use |
|---|---|
Coding | Feature work, debugging, repository changes and coding-agent workflows |
Professional work | Documents, analysis and multi-tool business processes |
Computer use | Supported GUI interaction through the Responses API |
Long context | Large repositories and document sets within the 922K input limit |
Reasoning control | Six effort levels from none through max |
Prompt caching | Reusable prefixes, diagnostics and explicit breakpoints |
The strongest documented fit is complex coding and agentic workflows at lower cost than GPT-6 Astra. A capable base model can still fail when retrieval is weak, tools are over-permissioned, instructions conflict or the system lacks verification. Evaluate the full application rather than the model in isolation.
Benchmark evidence
OpenAI reports 68.8% on DeepSWE v1.1 at maximum effort and 33.2% on AutomationBench at xhigh effort. It also reports that Sol approaches Claude Fable 5.1 on selected coding and workflow tests at substantially lower cost. These comparisons are provider-reported and use specific harnesses.
Evaluation | Reported result | Evidence boundary |
|---|---|---|
DeepSWE v1.1 | 68.8% at max effort | OpenAI-reported software-engineering result |
AutomationBench 1.0.6 | 33.2% at xhigh effort | OpenAI-reported multi-tool business result |
Agents’ Last Exam | 56.4% at max effort | OpenAI-reported long-horizon professional-work result |
OSWorld 2.0 offline | 60.5% at xhigh effort | OpenAI-reported computer-use result |
Internal factuality | About half as many mistakes as GPT-5.6 Sol | Difficult conversations previously flagged for errors |
Benchmark scores depend on model version, reasoning effort, sampling, tools, scaffolding and the exact dataset release. Do not combine scores from different harnesses into a synthetic ranking. Use public results to choose candidates, then test those candidates on frozen tasks from the intended workflow.
Artificial Analysis Intelligence Index
Artificial Analysis independently measures model intelligence, speed and cost. GPT-6 Sol scores 48 on the Artificial Analysis Intelligence Index at max effort. Replaced by GPT-6.1 Sol seven days after launch.
Metric | Value |
|---|---|
Intelligence Index score | 48 |
Effort setting | max |
Source | Artificial Analysis (independent) |
Last verified | October 2, 2026 |
The Artificial Analysis Intelligence Index combines scores across mathematics, reasoning, coding, instruction following and language tasks. The scale is not a percentage: the index reflects relative position across the models they track, not a share of correct answers. Compare scores only within the same effort setting and the same index version.
Pricing and access
GPT-6 Sol is available through the OpenAI API, ChatGPT Work and Codex for eligible users. It supports the Responses API and Chat Completions, although Chat Completions function calling requires reasoning effort set to none.
GPT-6.1 Sol is the newer Sol generation and should be the default comparison for new deployments. Keep GPT-6 Sol when a validated application depends on its exact behavior, then plan migration using a frozen regression suite.
Usage or access item | Current value |
|---|---|
Standard input | $2 per million tokens |
Cached input | $0.20 per million tokens |
Cache writes | $2.50 per million tokens |
Standard output | $10 per million tokens |
Batch and Flex | 50% below Standard rates |
Fast mode | 2 times Standard rates |
Prompts above 272K | 2 times input and cache; 1.5 times output for the full request |
Token price is only one part of total cost. Include cache behavior, long-context multipliers, tool calls, retries, wall time, failed-task recovery and human review. The useful comparison is cost per accepted result, not price per million tokens in isolation.
Limitations and deployment risks
- GPT-6.1 Sol is newer, more capable on OpenAI’s reported evaluations and cheaper for cached input.
- Prompts above 272,000 input tokens trigger higher rates across the full request.
- The strongest results use high or maximum reasoning effort and may have materially higher task cost.
- Chat Completions tool use is more limited than the Responses API path.
- Long context does not guarantee reliable retrieval or adherence across a million-token prompt.
- Computer-use and shell workflows need approval gates and rollback mechanisms.
- Migration can change output style, tool selection and caching behavior even when the API shape is similar.
High-stakes legal, medical, financial, scientific and security work needs qualified review. Log the exact model version and configuration, separate untrusted content from system instructions, restrict tools to the minimum required scope and require approval before irreversible or externally visible actions.
How to evaluate GPT-6 Sol
Build a frozen evaluation set of 20 to 50 real tasks. Include routine work, difficult edge cases, adversarial inputs, long-context examples and cases where the correct behavior is to stop or escalate. Compare GPT-6 Sol with GPT-5.6 Sol and at least one neighboring model using equivalent tools and source material.
Test area | What to record |
|---|---|
Task completion | Pass or fail against a written acceptance rubric |
Reliability | Repeated-run success rate, variance and silent failures |
Factuality | Unsupported claims, source use, quotations and citation accuracy |
Tool use | Wrong calls, retries, recovery and permission-boundary failures |
Long context | Retrieval accuracy, instruction retention and cost at realistic lengths |
Efficiency | Wall time, input, output, cached tokens, tool charges and review time |
Safety | Prompt injection, sensitive-data handling and irreversible-action controls |
Choose the least expensive configuration that clears the acceptance threshold with a safety margin. Re-run the suite when the model, system prompt, effort level, retrieval layer, tool definitions or approval policy changes.
Who should use it?
Situation | Recommendation |
|---|---|
Strong fit | complex coding and agentic workflows at lower cost than GPT-6 Astra |
Pilot first | Long-running agents, large contexts, computer use and workflows with several tools |
Escalate | Ambiguous or consequential work that does not reliably clear the evaluation threshold |
Avoid unsupervised use | Irreversible actions, sensitive data or high-stakes decisions without monitoring and approval |
A migration should be driven by measured outcomes. Keep GPT-5.6 Sol available during the pilot, record where each model succeeds or fails, and use routing when different task classes have different quality and cost requirements.
Frequently asked questions
Is GPT-6 Sol open source?
No. GPT-6 Sol is proprietary. Access, serving behavior and lifecycle decisions are controlled by OpenAI and supported distribution partners.
How much does GPT-6 Sol cost?
$2 per million tokens. Review the full pricing table above because cached input, long context, processing mode and platform can materially change the total.
What is GPT-6 Sol best used for?
Its strongest documented fit is complex coding and agentic workflows at lower cost than GPT-6 Astra. Start with a supervised pilot and retain human sign-off for consequential work.
Should I migrate from GPT-5.6 Sol?
Only after a side-by-side evaluation. Measure accepted-task quality, total cost, latency, output style, tool reliability and migration engineering. A newer model is not automatically the better operational choice.
How should benchmark evidence for GPT-6 Sol be interpreted?
Use provider results to identify promising workloads, then reproduce the comparison with the exact model configuration, tools and acceptance criteria that matter to the deployment. Do not infer a site ranking from benchmark claims gathered under a different harness.
Related model guides
Browse the Best AI Models directory.
Compare with the GPT-6 Astra guide.
Compare with the GPT 5.6 Sol guide.
Compare with the GPT 5.5 guide.
Official sources and update policy
OpenAI’s GPT-6 Sol and Luna announcement.
GPT-6 Sol API model documentation.
Checked October 2, 2026. We update this guide when OpenAI changes the specification, pricing, access, safety documentation or model lifecycle. Vendor benchmarks are attributed and are not presented as independent testing. No vendor payment or affiliate relationship determined inclusion.