Muse Spark 1.3: Complete Guide, Pricing, Specs and Use Cases

An independent guide to Muse Spark 1.3, including verified specifications, pricing, access, capabilities, limitations and a practical evaluation framework.

Follow in Google Search

What is Muse Spark 1.3?

Muse Spark 1.3 is Meta’s proprietary hosted model model for long-horizon agentic workflows, coding, multimodal analysis and tool-assisted professional work. It was released September 3, 2026. This guide separates documented specifications from vendor benchmark claims and gives teams a practical way to decide whether the model belongs in a real evaluation.

QUICK VERDICT: Muse Spark 1.3 is Meta’s frontier model and ranks third in our verified leaderboard. Its price and million-token context make it a candidate, but Meta’s claims still require workload-specific testing.

Field

Verified value

Provider

Meta

Release date

September 3, 2026

Availability

Available through Muse Code and Meta Model API; max reasoning was still pending additional safety testing at launch

License

Proprietary hosted model

Context window

1 million tokens

Maximum output

Endpoint-specific within the documented context budget; verify the active Meta Model API limit

Modalities

Text, image, video and document input; text and tool-directed output

API pricing

Standard endpoint: $1.25 input, $0.15 cached input and $4.25 output per million tokens; contributor endpoint: $0.10 input, $0.002 cached input and $0.20 output

Access

Muse Code and Meta Model API

Best fit

long-horizon agentic workflows, coding, multimodal analysis and tool-assisted professional work

Last verified

September 3, 2026

What changed with Muse Spark 1.3

Muse Spark 1.3 should not be evaluated as a name change. The important differences are the model’s reasoning controls, tool behavior, context policy, access route and economic profile. Those details determine whether a benchmark result transfers to production.

Area

What changed or matters

Agentic workflows

Better context tracking, plan repair, clarification and handling of multiple workflows in one long thread

Coding efficiency

Meta reports about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2

Instruction following

Improved preservation of detailed requirements across multi-step work

Safety

Stronger prompt-injection resistance and better recognition of irreversible actions

Reasoning modes

Existing modes launched immediately; max reasoning was scheduled after additional safety testing

The correct comparison baseline is Muse Spark 1.2. Teams planning a new deployment should also include Muse Spark 1.3 where access allows. Testing only the newest model or only the incumbent hides the migration cost, output-style changes and tool-use regressions that often matter more than a small benchmark gap.

Muse Spark 1.3 capabilities

The clearest fit is long-horizon agentic workflows, coding, multimodal analysis and tool-assisted professional work. That does not mean every task in those categories should use the model. A production system combines the base model with prompts, retrieval, tools, permissions, memory, retry logic and human review. The model card describes only one layer of that system.

Capability

Practical implication

Long context

1 million tokens. Test retrieval accuracy at realistic lengths rather than assuming every token receives equal attention.

Output capacity

Endpoint-specific within the documented context budget; verify the active Meta Model API limit. Long output is useful only when the verification process can keep up.

Modalities

Text, image, video and document input; text and tool-directed output. Confirm format-specific accuracy with your own files.

Agent use

Use explicit tool schemas, narrow permissions, approval gates and recoverable operations.

Reasoning

Record the exact effort level because cost, latency and quality can change materially.

For coding, evaluate repository navigation, test creation, regression rate, review burden and recovery after a failed tool call. For research, measure citation correctness, source coverage and whether the model distinguishes evidence from inference. For computer or browser use, record every unsafe click, wrong field, lost state and unapproved action.

Benchmarks and evidence

The evidence below comes from Meta or named launch partners. It is useful for identifying intended strengths, but it is not equivalent to an independent head-to-head test. Prompts, tools, reasoning budgets, sampling, infrastructure and scoring rules can differ.

Evidence

Reported result

How to interpret it

Current Overall ranking

81.6 and number three in The AI Leaderboard’s September 3 top eight

LiveBench-derived comparative score across seven evaluation categories, not a production success rate

Coding efficiency

About 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in Meta engineer comparisons

Vendor-reported internal comparison

Long-horizon behavior

Meta reports better plan correction, context tracking, clarification and user collaboration

Provider capability claim that depends on the agent harness

Evaluation coverage

Meta documents software engineering, terminal and multimodal evaluation methods

First-party methodology; inspect prompts, tools and scoring before comparison

Treat benchmark results as a shortlist signal. Before purchasing or migrating, reproduce representative work under one frozen protocol. Keep model snapshots, reasoning effort, tool access and token budgets constant. Report repeated-run pass rates and total cost per accepted result, not a single best attempt.

Pricing, access and deployment

Muse Code and Meta Model API. The current pricing reference is Standard endpoint: $1.25 input, $0.15 cached input and $4.25 output per million tokens; contributor endpoint: $0.10 input, $0.002 cached input and $0.20 output. Provider pricing changes frequently and can include cache rates, batch discounts, regional premiums, long-context multipliers, priority processing and tool-call fees. Recheck the official pricing page before budgeting.

Cost driver

What to measure

Input tokens

Prompt, retrieved context, tool results and repeated history

Output tokens

Visible answer plus any billable reasoning or generated artifacts

Caching

Eligible repeated prefixes, cache-read price and expiration policy

Tools

Search, computer use, code execution and third-party API fees

Retries

Failed runs, verifier loops and human rework

Success-adjusted cost

Total spend divided by deliverables that pass review

A cheaper token price can lose to a more expensive model if it takes more steps, retries more often or produces work that needs heavy correction. Conversely, a frontier model can be wasteful when a smaller model already passes the task rubric. Route by measured task difficulty rather than brand prestige.

Muse Spark 1.3 limitations

  • Muse Spark 1.3 is proprietary and available only through Meta-controlled access routes.
  • Max reasoning was not available at launch and was pending additional safety testing.
  • Meta’s efficiency and capability comparisons are vendor-reported and need independent reproduction.
  • A one-million-token context window does not guarantee reliable retrieval or instruction retention across the full window.
  • Tool use, browser access and long-running coding workflows require least privilege, approval gates and recoverable operations.

High-stakes medical, legal, financial, security and scientific work requires qualified review. Store the exact model identifier, prompt version, tools, source documents and approvals for each consequential run. Build a rollback path before granting write access to repositories, browsers, databases or cloud infrastructure.

Risk

Minimum control

Hallucination

Require source checks or executable tests

Prompt injection

Separate untrusted content from instructions and restrict tools

Over-permission

Use least privilege and approval gates

Silent model change

Pin snapshots where possible and run regression tests

Data exposure

Review retention, regional processing and provider terms

Runaway cost

Set token, time, tool-call and retry budgets

How to evaluate Muse Spark 1.3

Create 20 to 50 tasks from real work. Freeze the tasks and rubric before testing. Include easy tasks, normal tasks, edge cases and adversarial inputs. Give each model equivalent tools and enough budget to finish, but cap time and retries. Repeat non-deterministic runs so one lucky result does not decide the winner.

Test area

Record

Task completion

Pass or fail plus rubric score

Reliability

Repeated-run success and variance

Quality

Factuality, instruction adherence and usefulness

Tool use

Wrong calls, retries, recovery and permission errors

Efficiency

Wall time, tokens, cache use, tool fees and human review

Safety

Unsafe actions, injection response and sensitive-data handling

Migration

Prompt changes, integration work and output-style regressions

Compare Muse Spark 1.3 with Muse Spark 1.2 and Muse Spark 1.3. Choose the least expensive configuration that meets the acceptance threshold with an adequate safety margin. Re-run the suite after any model snapshot, system prompt, retrieval or tool change.

Frequently asked questions

Is Muse Spark 1.3 available now?

Available through Muse Code and Meta Model API; max reasoning was still pending additional safety testing at launch. The documented access routes are Muse Code and Meta Model API. Availability can differ by region, plan and partner platform.

Is Muse Spark 1.3 open source?

No. Muse Spark 1.3 is proprietary. Access is controlled by Meta and supported distribution partners.

How much does Muse Spark 1.3 cost?

Standard endpoint: $1.25 input, $0.15 cached input and $4.25 output per million tokens; contributor endpoint: $0.10 input, $0.002 cached input and $0.20 output. Budget with measured end-to-end workloads because token rates alone omit retries, tools, caching, long-context multipliers and review time.

What is Muse Spark 1.3 best used for?

Its strongest documented fit is long-horizon agentic workflows, coding, multimodal analysis and tool-assisted professional work. Start with a supervised pilot and keep human sign-off for consequential work.

Should I migrate from Muse Spark 1.2?

Only after a side-by-side evaluation. Measure pass rate, total cost, latency, output style, tool reliability and the engineering work required to migrate. A newer model is not automatically the better operational choice.

Related model guides

Start with the broader Best AI Models directory.

Compare with the Muse Spark 1.1 guide.

Compare with the Meta models hub.

Official sources and update policy

Meta Muse Spark 1.3 announcement.

Muse Spark 1.3 evaluation methodology.

Muse Spark model page.

Meta Model API access and pricing.

Checked September 3, 2026. We update this guide when the provider changes the model specification, pricing, access, licensing or safety documentation. Vendor benchmarks are attributed and are not presented as independent testing. No vendor payment or affiliate relationship determined inclusion.

Author

Dr. Elena Vasquez

PhD in Computer Science, Stanford University (2018); MS in Machine Learning, Carnegie Mellon University. Research on scaling laws, evaluation methodologies, and robustness in large neural models.