# Meta's Muse Spark 1.1 Is Not A Catch-Up Model. It Is A Paid Agent Platform

**Plutonous** | July 12, 2026 | 10 min read

> Meta Muse Spark 1.1 API pricing, 1M-token context, benchmarks, coding agents, and what Meta's paid agent platform means for developers.

Tags: Meta, Muse Spark 1.1, AI Agents, Model APIs, Artificial Analysis, AI Pricing, Long Context, AI Safety

---

**TL;DR:** Meta launched Muse Spark 1.1 and its public-preview Model API on July 9, pricing the reasoning model at **$1.25** per million input tokens and **$4.25** per million output tokens, with a **1 million-token** context window.<sup><a href="#source-1">[1]</a></sup><sup><a href="#source-2">[2]</a></sup> Artificial Analysis scores the `xhigh` configuration at **51** on its Intelligence Index, up **8 points** from Muse Spark 1.0, at an estimated **$0.26** per standardized Index task.<sup><a href="#source-3">[3]</a></sup> The real story isn't Meta winning a leaderboard. It is Meta turning its distribution advantage into a paid, long-context agent platform.

Meta has spent years proving it can distribute AI to billions of people. That was never the difficult business problem. The difficult problem was turning that reach into a developer platform that enterprises trust with code, tools, documents, and money.

Muse Spark 1.1 is Meta's first serious answer. The closed-weight model powers Meta AI's Thinking mode and is available through the new Meta Model API public preview.<sup><a href="#source-1">[1]</a></sup> Meta says it is multimodal and built for tool use, computer use, coding, and orchestration. Artificial Analysis' evaluated endpoint currently lists text and image input with text output, a useful reminder that product claims and a tested API configuration are not interchangeable.<sup><a href="#source-2">[2]</a></sup>

That distinction matters. A cheap chat model is not an agent platform. A strong benchmark score is not an agent platform either. Long-lived context, cache economics, function calling, safety controls, and an API developers can actually ship are the platform. Meta is finally selling that whole bundle.


### Why This Matters Now

**Meta's break:** Muse Spark 1.1 shifts Meta from AI distribution to paid AI infrastructure.<br/>
**The evidence:** Artificial Analysis puts the `xhigh` configuration at 51, a genuine eight-point improvement, but still below the 60-point leader in its July 12 snapshot.<sup><a href="#source-3">[3]</a></sup><br/>
**The bet:** Meta is pricing a million-token, agent-ready model cheaply enough to make persistent workflows practical, then using its consumer footprint to create demand above the API.


/images/articles/muse-spark-1-1-agent-runtime/muse-spark-1-1-agent-runtime-cover.webp

Editorial engraving of an ink-black orchestration engine linking crimson threads to abstract code, document, browser, and data workstations on a cream newsprint field.

Conceptual illustration: Meta's commercial wager is not simply a higher model score. It is an agent runtime that holds context, routes work across tools, and makes long-running tasks economically viable.

1672

941

16/9

cover

940


## The Benchmark Reality: Good Enough To Matter, Not Good Enough To Declare Victory

The screenshot-friendly number is **51**. On Artificial Analysis' current Intelligence Index, Muse Spark 1.1 sits in a crowded frontier cluster, 3 points behind [Grok 4.5](/news/grok-45-frontier-model-rankings-agent-economics) at 54 and 9 behind Claude Fable 5 at 60 in the July 12 view.<sup><a href="#source-2">[2]</a></sup> That is a formidable result for a three-month iteration. It is not a clean frontier takeover.

Artificial Analysis says it supported Meta with pre-release evaluation. Its measurement is more useful than a vendor chart, but not entirely arm's-length in the strictest sense.<sup><a href="#source-3">[3]</a></sup> Its composite uses nine standardized evaluations across agents, coding, scientific reasoning, general knowledge, and long-context work. That makes it decision-relevant, not definitive.<sup><a href="#source-4">[4]</a></sup>

The gains are specific. Artificial Analysis records a **+12-point** Coding Index increase, SciCode rising from **52% to 58%**, Humanity's Last Exam rising from **40% to 45%**, and GDPval-AA v2 climbing **232 Elo** points from 1,144 to 1,376.<sup><a href="#source-3">[3]</a></sup> SciCode is the standout, ranking the model third in Artificial Analysis' launch comparison at 58%.

What's often overlooked is the reliability tradeoff. Its AA-Omniscience score rose from **4 to 18** largely because the model attempted fewer questions. Hallucination fell from **73% to 38%**, while reported accuracy slipped from **45% to 41%** and attempt rate fell from **95% to 82%**.<sup><a href="#source-3">[3]</a></sup> That is a sensible production change. A model that knows when not to improvise is safer in an agent loop. It is not the same thing as a model that suddenly knows more.


### Muse Spark 1.1: What The Numbers Actually Say
- Third-party measurement
- Meta's published claim
- What a buyer should conclude

- feature: Overall capability; values: - 51 AA Intelligence Index
- Broad improvement across capability benchmarks
- Competitive frontier-cluster model, not the category leader
- feature: Coding; values: - 58% SciCode; #3 in AA launch comparison
- Strong coding and agentic affordances
- Worth testing for coding agents, especially where cost matters
- feature: Knowledge reliability; values: - Lower hallucination, but lower attempt rate and accuracy
- Stronger robustness and behavior results
- Evaluate task accuracy and abstention policy in your own harness
- feature: Safety; values: - No independent full safety replication at launch
- Residual risk is moderate or lower after mitigations
- Use allowlists, sandboxing, and audit trails


51

Artificial Analysis Intelligence Index score for Muse Spark 1.1 (xhigh)

An eight-point gain from Muse Spark 1.0, but not proof of universal frontier parity.


/images/articles/muse-spark-1-1-agent-runtime/muse-spark-1-1-benchmark-reality.webp

Eight abstract ink-black AI engine forms aligned beneath a shared measurement line, with one crimson marker only slightly above the tightly grouped field.

Artificial Analysis places Muse Spark 1.1 in a close competitive cluster on its composite index. This editorial illustration represents relative proximity, not a literal benchmark chart or a decisive leader.

1672

941

16/9

cover

940


## The Economics: Meta Is Discounting The Agent Loop, Not Just The Token

At face value, Muse Spark 1.1 costs **$1.25/$4.25** per million input and output tokens. Cache hits cost **$0.15** per million input tokens, an **88%** discount from the standard input price.<sup><a href="#source-2">[2]</a></sup> New Meta Model API accounts receive **$20** in launch credits, according to Meta and Reuters.<sup><a href="#source-1">[1]</a></sup><sup><a href="#source-5">[5]</a></sup>

The more important metric is not the price card. It is the cost of finishing work. Artificial Analysis estimates **$0.26** per Intelligence Index task, compared with $0.37 for GLM-5.2 and $0.89 for GPT-5.4 in its setup.<sup><a href="#source-3">[3]</a></sup> That is not a promise about your support bot, coding agent, or research workflow. Its calculation includes a standardized workload, token consumption, cache assumptions, and the model's own reasoning allocation.

The uncomfortable truth is that cheap output pricing can hide a large bill. The `xhigh` configuration generated **94 million** output tokens over the full Artificial Analysis index, against a **60 million** median for comparable models.<sup><a href="#source-2">[2]</a></sup> Teams should meter completed-task cost, tool retries, and completion length, not only the headline input price.


### Muse Spark 1.1: The Operating Numbers
Artificial Analysis measurements for the `xhigh` configuration on Meta's first-party API, alongside Meta's stated pricing.

- label: Input price; value: $1.25; description: Per 1 million tokens.; trendText: low list price
- label: Output price; value: $4.25; description: Per 1 million tokens.; trendText: agent-loop leverage
- label: Output speed; value: 116.3 t/s; description: AA measurement with 1.05-second time to first token.; trendText: configuration-specific
- label: Context window; value: 1M; description: Tokens, up from 262,000 in Spark 1.0.; trendText: 4x larger
- label: Indexed task cost; value: $0.26; description: AA estimate, not a universal production-task price.; trendText: workload dependent

Sources: Meta launch materials and Artificial Analysis. Throughput and task-cost figures are configuration- and workload-specific.

source


/images/articles/muse-spark-1-1-agent-runtime/muse-spark-1-1-agent-economics.webp

A black mechanical computing engine recirculates task tokens through a crimson cache loop, with a separate metered compute tower at right.

Illustration: cache reuse can reduce the repeat-compute burden of long-running agent workflows; it does not depict Meta's infrastructure or quantify a specific saving.

1672

941

16/9

cover

940


Here's the genius: Meta is not merely racing to the lowest token price. It is trying to make the expensive agent loop economically ordinary. Give an agent a codebase, policy manual, document archive, live search tool, and structured-output requirement. The winning model is the one that can keep that state in play without turning every tool iteration into a premium purchase. That is the same economic logic behind the fight to own the [coding-model loop](/news/spacex-cursor-60b-composer-model-race), not a marginal API-price tweak.

## One Million Tokens Changes The Product Shape

Meta expanded Muse Spark's context window from **262,000** tokens to **1 million**.<sup><a href="#source-2">[2]</a></sup> The usual long-context pitch is “upload more documents.” The strategic use is more demanding: carry policy, prior work, user instructions, tool results, repository state, and a multi-step plan in one working memory.


### From Chat Prompt To Persistent Agent
A million-token window changes which parts of an agent workflow can remain in one model session.

- title: Load the operating context; description: Bring in the repository, policy corpus, active tickets, schemas, and user constraints.; volume: Up to 1M tokens; time: Session start
- title: Reason across dependencies; description: Compare instructions against source material and produce a structured plan before touching tools.; volume: Agent planning; time: Before action
- title: Call approved tools; description: Meta exposes tool and function calling. The application decides which tools exist and what actions they may take.; volume: Bounded execution; time: Agent loop
- title: Keep evidence with the answer; description: Require tests, citations, and an audit trail before consequential changes.; volume: Evidence and logs; time: Before handoff


/images/articles/muse-spark-1-1-agent-runtime/muse-spark-1-1-million-token-context.webp

An ink-black central memory node linked by a crimson thread to a vast field of archival pages and document paths.

An editorial illustration of the long-context and memory-continuity thesis. It is conceptual, not a depiction of Meta's product interface or a claim about its architecture.

1672

941

16/9

cover

940


The real story isn't that retrieval engineering disappears. It gets more important. More context can mean more stale instructions, hostile content, irrelevant evidence, and prompt-injection opportunities. [Recursive language models](/news/recursive-language-models-context-window-breakthrough) make the same point from the opposite direction: useful long-context systems need deliberate external state and evidence handling, not blind prompt accumulation. Meta's own report recommends tool allowlists and workspace isolation for API deployments.<sup><a href="#source-6">[6]</a></sup> A huge context window without access controls is not an agent architecture. It is a larger attack surface.

## The API Is The Real Release: Meta Wants To Sit Above The Open-Weight Layer

For years, Meta's AI identity was synonymous with Llama and open-weight distribution. Muse Spark 1.1 is proprietary. Meta has not disclosed a parameter count.<sup><a href="#source-2">[2]</a></sup> That does not mean Meta abandoned open models. It means the company is separating its research-distribution strategy from its highest-value product strategy.

The Model API is a public preview, has OpenAI-compatible tooling, and exposes structured output, parallel tool calling, MCP/custom-skill use, and built-in web search according to Meta's developer materials.<sup><a href="#source-1">[1]</a></sup><sup><a href="#source-7">[7]</a></sup> This is the flywheel: consumer distribution supplies use cases and feedback; the API monetizes developers building the workflows consumers and businesses will expect.


/images/articles/muse-spark-1-1-agent-runtime/muse-spark-1-1-api-shift.webp

Editorial illustration of open model artifacts feeding through a crimson metering valve into a controlled black agent-runtime gateway.

A conceptual rendering of the strategic shift described in this article: turning openly distributed model artifacts into a controlled, metered agent-service layer. It is not a depiction of Meta infrastructure.

1672

941

16/9

cover

940


While competitors sell model families and enterprise contracts, Meta is attempting to sell the missing layer between them: context-rich agents that can appear in the apps people already use. The company has a better consumer distribution channel than almost any frontier lab. It now needs a reason for developers to keep complex work on its API instead of routing through existing defaults.

## Safety Claims Need Application Controls, Not Applause

Meta's evaluation report says that, without mitigations, the company cannot rule out that Muse Spark 1.1 meets its “high risk” capability threshold in chemical and biological and cybersecurity domains. After deployment mitigations, Meta classifies residual risk as “moderate or lower.”<sup><a href="#source-6">[6]</a></sup>

That is not an argument against release. It is an argument against magical thinking about agent safety. As models get better at code, tools, and long-horizon tasks, the question is whether the application prevents a capable system from receiving broad credentials, executing uncontrolled actions, or treating retrieved content as authority. [Loop engineering](/news/loop-engineering-designing-agent-stop-conditions) is the practical companion: explicit stop conditions and independent verification are product decisions, not model settings.


A model card can describe the guardrails Meta ships. It cannot define the permissions your application hands to the model.

The operational lesson of Muse Spark 1.1


### The Key Risk Is Permission, Not Just Hallucination

Muse Spark 1.1 is designed for agentic affordances. Start with read-only tools, narrow scopes, confirmation gates for consequential changes, isolated workspaces, and complete action logs. A million-token context window should make your evidence trail better, not your blast radius larger.


/images/articles/muse-spark-1-1-agent-runtime/muse-spark-1-1-permission-controls.webp

An abstract black robotic hand halted at a crimson permission barrier before a set of contained geometric tool forms.

A conceptual illustration of constrained agent operation: capability only becomes useful when access to tools is explicitly bounded and authorized.

1536

1024

3/2

cover

940


## Muse Spark 1.1 API: Pricing, Context, Comparisons And Access

### How much does the Muse Spark 1.1 API cost?

Meta prices Muse Spark 1.1 at **$1.25 per million input tokens** and **$4.25 per million output tokens**. Cached input costs **$0.15 per million tokens**, while every new Meta Model API account starts with **$20** in one-time credits.<sup><a href="#source-1">[1]</a></sup><sup><a href="#source-5">[5]</a></sup> Those rates are the list price, not a guarantee of completed-task cost. A long-running agent can spend far more through output volume, retries, and tool loops.

### Does Muse Spark 1.1 have a 1 million-token context window?

Yes. Artificial Analysis lists a **1 million-token** context window for the evaluated `xhigh` configuration, up from **262,000** in Muse Spark 1.0.<sup><a href="#source-2">[2]</a></sup> The useful question is not how many documents fit. It is whether the agent can keep the right evidence, instructions, and tool state salient during a real workflow.

### Is Muse Spark 1.1 open source or a Llama model?

No. Muse Spark 1.1 is a **proprietary, closed-weight** model from Meta Superintelligence Labs. Meta has not disclosed a parameter count.<sup><a href="#source-2">[2]</a></sup> Llama remains Meta's open-weight strategy. Muse Spark is Meta's paid agent-service strategy.

### Is Muse Spark 1.1 OpenAI-compatible?

Meta's developer materials say Meta Model API is compatible with OpenAI SDK tooling and common OpenAI-compatible agent setups.<sup><a href="#source-7">[7]</a></sup> Compatibility reduces integration friction. It does not make model behavior, output format edge cases, rate limits, or safety controls identical, so teams should test the exact workflow they plan to ship.

### What are Muse Spark API rate limits and access rules?

Muse Spark 1.1 is in Meta Model API public preview. Meta's rate-limit details are account-specific and shown in the authenticated developer documentation, so an article should not freeze a temporary quota into a supposedly universal specification.<sup><a href="#source-9">[9]</a></sup> Check the current dashboard before committing a production workload, and use the $20 launch credit for a constrained evaluation first.

### Muse Spark 1.1 vs Claude and GPT: what does the benchmark snapshot say?

In Artificial Analysis' July 12 snapshot, Muse Spark 1.1 scored **51**, versus **60** for Claude Fable 5 and **59** for GPT-5.6 Sol.<sup><a href="#source-2">[2]</a></sup> That does not settle every workload. The configurations, reasoning effort, task mix, cost assumptions, and safety policies differ. It does establish the commercial question: can Meta's lower price and one-million-token context overcome a capability gap on the specific agent task a buyer actually has?

## The Real Contest: Can Meta Own The Work Between Prompt And Action?

Muse Spark 1.1 is not the model that makes every other model irrelevant. Its 51-point third-party score says the opposite: capability is clustering tightly enough that raw intelligence alone is no longer a defensible product category.<sup><a href="#source-3">[3]</a></sup>

That is precisely why Meta's move matters. It has combined competitive capability, first-party serving, a one-million-token context window, low cache pricing, tool calling, and a consumer distribution machine. No single ingredient is exclusive. The integration is the wager.

The winners will not be the labs that post the largest benchmark number for one week. They will be the platforms that make agents cheap to operate, safe to constrain, easy to evaluate, and native to real work. Muse Spark 1.1 gives Meta a credible seat at that table. The next release has to prove that Meta can own the workflow, not just enter the chart.


### What To Do With Muse Spark 1.1
- Treat the 51 Intelligence Index score as a reason to test, not a reason to standardize. The result is strong but not category-leading.
- Benchmark completed-task cost, not only token price. The low rates and cache discount are compelling, while reasoning output can be verbose.
- Use the million-token window for evidence-rich, bounded workflows. Do not use it to bypass retrieval hygiene or prompt-injection defenses.
- Keep agent permissions narrow. Tool allowlists and workspace isolation should be the baseline for production deployment.
- Watch whether Meta turns API usage into distribution inside consumer products. That integration, not one benchmark, is the strategic prize.


## Sources & References

<a id="source-1"></a>
1. [Introducing Muse Spark 1.1](https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/)

<a id="source-2"></a>
2. [Muse Spark 1.1 (xhigh): Intelligence, Performance & Price](https://artificialanalysis.ai/models/muse-spark-1-1/)

<a id="source-3"></a>
3. [Muse Spark 1.1: Meta gains 8 Intelligence Index points](https://artificialanalysis.ai/articles/muse-spark-1-1-everything-you-need-to-know/)

<a id="source-4"></a>
4. [Artificial Analysis Intelligence Benchmarking Methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking)

<a id="source-5"></a>
5. [Meta jumps into AI coding market to chase Anthropic and OpenAI](https://www.cnbc.com/2026/07/09/meta-jumps-into-ai-coding-market-to-chase-anthropic-and-openai.html)

<a id="source-6"></a>
6. [Muse Spark 1.1 Evaluation Report](https://ai.meta.com/static-resource/muse-spark-1-1-evaluation-report)

<a id="source-7"></a>
7. [Build with Muse Spark on Meta Model API](https://developer.meta.com/ai/resources/blog/build-with-muse-spark/)

<a id="source-8"></a>
8. [Meta enters the crowded AI coding battle with Muse Spark 1.1](https://techcrunch.com/2026/07/09/meta-enters-the-crowded-ai-coding-battle-with-muse-spark-1-1/)

<a id="source-9"></a>
9. [Meta Model API Pricing And Rate Limits](https://dev.meta.ai/docs/getting-started/pricing-rate-limits)


*Last updated: July 13, 2026*

---

*Source: [LLM Rumors](https://www.llmrumors.com/news/meta-muse-spark-11-paid-agent-platform)*
