Back to News
Meta

Meta's Muse Spark 1.1 Is Not A Catch-Up Model. It Is A Paid Agent Platform

LLM Rumors··10 min read·...
MetaMuse Spark 1.1AI AgentsModel APIsArtificial AnalysisAI PricingLong ContextAI Safety
Meta's Muse Spark 1.1 Is Not A Catch-Up Model. It Is A Paid Agent Platform

TL;DR: Meta launched Muse Spark 1.1 and its public-preview Model API on July 9, pricing the reasoning model at $1.25 per million input tokens and $4.25 per million output tokens, with a 1 million-token context window.[1][2] Artificial Analysis scores the xhigh configuration at 51 on its Intelligence Index, up 8 points from Muse Spark 1.0, at an estimated $0.26 per standardized Index task.[3] The real story isn't Meta winning a leaderboard. It is Meta turning its distribution advantage into a paid, long-context agent platform.

Meta has spent years proving it can distribute AI to billions of people. That was never the difficult business problem. The difficult problem was turning that reach into a developer platform that enterprises trust with code, tools, documents, and money.

Muse Spark 1.1 is Meta's first serious answer. The closed-weight model powers Meta AI's Thinking mode and is available through the new Meta Model API public preview.[1] Meta says it is multimodal and built for tool use, computer use, coding, and orchestration. Artificial Analysis' evaluated endpoint currently lists text and image input with text output, a useful reminder that product claims and a tested API configuration are not interchangeable.[2]

That distinction matters. A cheap chat model is not an agent platform. A strong benchmark score is not an agent platform either. Long-lived context, cache economics, function calling, safety controls, and an API developers can actually ship are the platform. Meta is finally selling that whole bundle.

NOTE

Why This Matters Now

Meta's break: Muse Spark 1.1 shifts Meta from AI distribution to paid AI infrastructure.
The evidence: Artificial Analysis puts the xhigh configuration at 51, a genuine eight-point improvement, but still below the 60-point leader in its July 12 snapshot.[3]
The bet: Meta is pricing a million-token, agent-ready model cheaply enough to make persistent workflows practical, then using its consumer footprint to create demand above the API.

Editorial engraving of an ink-black orchestration engine linking crimson threads to abstract code, document, browser, and data workstations on a cream newsprint field.
Conceptual illustration: Meta's commercial wager is not simply a higher model score. It is an agent runtime that holds context, routes work across tools, and makes long-running tasks economically viable.

The Benchmark Reality: Good Enough To Matter, Not Good Enough To Declare Victory

The screenshot-friendly number is 51. On Artificial Analysis' current Intelligence Index, Muse Spark 1.1 sits in a crowded frontier cluster, 3 points behind Grok 4.5 at 54 and 9 behind Claude Fable 5 at 60 in the July 12 view.[2] That is a formidable result for a three-month iteration. It is not a clean frontier takeover.

Artificial Analysis says it supported Meta with pre-release evaluation. Its measurement is more useful than a vendor chart, but not entirely arm's-length in the strictest sense.[3] Its composite uses nine standardized evaluations across agents, coding, scientific reasoning, general knowledge, and long-context work. That makes it decision-relevant, not definitive.[4]

The gains are specific. Artificial Analysis records a +12-point Coding Index increase, SciCode rising from 52% to 58%, Humanity's Last Exam rising from 40% to 45%, and GDPval-AA v2 climbing 232 Elo points from 1,144 to 1,376.[3] SciCode is the standout, ranking the model third in Artificial Analysis' launch comparison at 58%.

What's often overlooked is the reliability tradeoff. Its AA-Omniscience score rose from 4 to 18 largely because the model attempted fewer questions. Hallucination fell from 73% to 38%, while reported accuracy slipped from 45% to 41% and attempt rate fell from 95% to 82%.[3] That is a sensible production change. A model that knows when not to improvise is safer in an agent loop. It is not the same thing as a model that suddenly knows more.

Muse Spark 1.1: What The Numbers Actually Say

FeatureThird-party measurementMeta's published claimWhat a buyer should conclude
Overall capability51 AA Intelligence IndexBroad improvement across capability benchmarksCompetitive frontier-cluster model, not the category leader
Coding58% SciCode; #3 in AA launch comparisonStrong coding and agentic affordancesWorth testing for coding agents, especially where cost matters
Knowledge reliabilityLower hallucination, but lower attempt rate and accuracyStronger robustness and behavior resultsEvaluate task accuracy and abstention policy in your own harness
SafetyNo independent full safety replication at launchResidual risk is moderate or lower after mitigationsUse allowlists, sandboxing, and audit trails
Loading interactive graphic
Loading interactive graphic
51
Artificial Analysis Intelligence Index score for Muse Spark 1.1 (xhigh)

An eight-point gain from Muse Spark 1.0, but not proof of universal frontier parity.

Eight abstract ink-black AI engine forms aligned beneath a shared measurement line, with one crimson marker only slightly above the tightly grouped field.
Artificial Analysis places Muse Spark 1.1 in a close competitive cluster on its composite index. This editorial illustration represents relative proximity, not a literal benchmark chart or a decisive leader.

The Economics: Meta Is Discounting The Agent Loop, Not Just The Token

At face value, Muse Spark 1.1 costs $1.25/$4.25 per million input and output tokens. Cache hits cost $0.15 per million input tokens, an 88% discount from the standard input price.[2] New Meta Model API accounts receive $20 in launch credits, according to Meta and Reuters.[1][5]

The more important metric is not the price card. It is the cost of finishing work. Artificial Analysis estimates $0.26 per Intelligence Index task, compared with $0.37 for GLM-5.2 and $0.89 for GPT-5.4 in its setup.[3] That is not a promise about your support bot, coding agent, or research workflow. Its calculation includes a standardized workload, token consumption, cache assumptions, and the model's own reasoning allocation.

The uncomfortable truth is that cheap output pricing can hide a large bill. The xhigh configuration generated 94 million output tokens over the full Artificial Analysis index, against a 60 million median for comparable models.[2] Teams should meter completed-task cost, tool retries, and completion length, not only the headline input price.

Muse Spark 1.1: The Operating Numbers

Artificial Analysis measurements for the `xhigh` configuration on Meta's first-party API, alongside Meta's stated pricing.

$1.25
Input price

Per 1 million tokens.

+ low list price
$4.25
Output price

Per 1 million tokens.

+ agent-loop leverage
116.3 t/s
Output speed

AA measurement with 1.05-second time to first token.

+ configuration-specific
1M
Context window

Tokens, up from 262,000 in Spark 1.0.

+ 4x larger
$0.26
Indexed task cost

AA estimate, not a universal production-task price.

= workload dependent
Sources: Meta launch materials and Artificial Analysis. Throughput and task-cost figures are configuration- and workload-specific.
A black mechanical computing engine recirculates task tokens through a crimson cache loop, with a separate metered compute tower at right.
Illustration: cache reuse can reduce the repeat-compute burden of long-running agent workflows; it does not depict Meta's infrastructure or quantify a specific saving.

Here's the genius: Meta is not merely racing to the lowest token price. It is trying to make the expensive agent loop economically ordinary. Give an agent a codebase, policy manual, document archive, live search tool, and structured-output requirement. The winning model is the one that can keep that state in play without turning every tool iteration into a premium purchase. That is the same economic logic behind the fight to own the coding-model loop, not a marginal API-price tweak.

One Million Tokens Changes The Product Shape

Meta expanded Muse Spark's context window from 262,000 tokens to 1 million.[2] The usual long-context pitch is “upload more documents.” The strategic use is more demanding: carry policy, prior work, user instructions, tool results, repository state, and a multi-step plan in one working memory.

From Chat Prompt To Persistent Agent

A million-token window changes which parts of an agent workflow can remain in one model session.

1

Load the operating context

Bring in the repository, policy corpus, active tickets, schemas, and user constraints.

Time:Session start
Scale:Up to 1M tokens
2

Reason across dependencies

Compare instructions against source material and produce a structured plan before touching tools.

Time:Before action
Scale:Agent planning
Key Step
3

Call approved tools

Meta exposes tool and function calling. The application decides which tools exist and what actions they may take.

Time:Agent loop
Scale:Bounded execution
4

Keep evidence with the answer

Require tests, citations, and an audit trail before consequential changes.

Time:Before handoff
Scale:Evidence and logs
An ink-black central memory node linked by a crimson thread to a vast field of archival pages and document paths.
An editorial illustration of the long-context and memory-continuity thesis. It is conceptual, not a depiction of Meta's product interface or a claim about its architecture.

The real story isn't that retrieval engineering disappears. It gets more important. More context can mean more stale instructions, hostile content, irrelevant evidence, and prompt-injection opportunities. Recursive language models make the same point from the opposite direction: useful long-context systems need deliberate external state and evidence handling, not blind prompt accumulation. Meta's own report recommends tool allowlists and workspace isolation for API deployments.[6] A huge context window without access controls is not an agent architecture. It is a larger attack surface.

The API Is The Real Release: Meta Wants To Sit Above The Open-Weight Layer

For years, Meta's AI identity was synonymous with Llama and open-weight distribution. Muse Spark 1.1 is proprietary. Meta has not disclosed a parameter count.[2] That does not mean Meta abandoned open models. It means the company is separating its research-distribution strategy from its highest-value product strategy.

The Model API is a public preview, has OpenAI-compatible tooling, and exposes structured output, parallel tool calling, MCP/custom-skill use, and built-in web search according to Meta's developer materials.[1][7] This is the flywheel: consumer distribution supplies use cases and feedback; the API monetizes developers building the workflows consumers and businesses will expect.

Editorial illustration of open model artifacts feeding through a crimson metering valve into a controlled black agent-runtime gateway.
A conceptual rendering of the strategic shift described in this article: turning openly distributed model artifacts into a controlled, metered agent-service layer. It is not a depiction of Meta infrastructure.

While competitors sell model families and enterprise contracts, Meta is attempting to sell the missing layer between them: context-rich agents that can appear in the apps people already use. The company has a better consumer distribution channel than almost any frontier lab. It now needs a reason for developers to keep complex work on its API instead of routing through existing defaults.

Safety Claims Need Application Controls, Not Applause

Meta's evaluation report says that, without mitigations, the company cannot rule out that Muse Spark 1.1 meets its “high risk” capability threshold in chemical and biological and cybersecurity domains. After deployment mitigations, Meta classifies residual risk as “moderate or lower.”[6]

That is not an argument against release. It is an argument against magical thinking about agent safety. As models get better at code, tools, and long-horizon tasks, the question is whether the application prevents a capable system from receiving broad credentials, executing uncontrolled actions, or treating retrieved content as authority. Loop engineering is the practical companion: explicit stop conditions and independent verification are product decisions, not model settings.

A model card can describe the guardrails Meta ships. It cannot define the permissions your application hands to the model.

The operational lesson of Muse Spark 1.1
WARNING

The Key Risk Is Permission, Not Just Hallucination

Muse Spark 1.1 is designed for agentic affordances. Start with read-only tools, narrow scopes, confirmation gates for consequential changes, isolated workspaces, and complete action logs. A million-token context window should make your evidence trail better, not your blast radius larger.

Loading interactive graphic
An abstract black robotic hand halted at a crimson permission barrier before a set of contained geometric tool forms.
A conceptual illustration of constrained agent operation: capability only becomes useful when access to tools is explicitly bounded and authorized.

Muse Spark 1.1 API: Pricing, Context, Comparisons And Access

How much does the Muse Spark 1.1 API cost?

Meta prices Muse Spark 1.1 at $1.25 per million input tokens and $4.25 per million output tokens. Cached input costs $0.15 per million tokens, while every new Meta Model API account starts with $20 in one-time credits.[1][5] Those rates are the list price, not a guarantee of completed-task cost. A long-running agent can spend far more through output volume, retries, and tool loops.

Does Muse Spark 1.1 have a 1 million-token context window?

Yes. Artificial Analysis lists a 1 million-token context window for the evaluated xhigh configuration, up from 262,000 in Muse Spark 1.0.[2] The useful question is not how many documents fit. It is whether the agent can keep the right evidence, instructions, and tool state salient during a real workflow.

Is Muse Spark 1.1 open source or a Llama model?

No. Muse Spark 1.1 is a proprietary, closed-weight model from Meta Superintelligence Labs. Meta has not disclosed a parameter count.[2] Llama remains Meta's open-weight strategy. Muse Spark is Meta's paid agent-service strategy.

Is Muse Spark 1.1 OpenAI-compatible?

Meta's developer materials say Meta Model API is compatible with OpenAI SDK tooling and common OpenAI-compatible agent setups.[7] Compatibility reduces integration friction. It does not make model behavior, output format edge cases, rate limits, or safety controls identical, so teams should test the exact workflow they plan to ship.

What are Muse Spark API rate limits and access rules?

Muse Spark 1.1 is in Meta Model API public preview. Meta's rate-limit details are account-specific and shown in the authenticated developer documentation, so an article should not freeze a temporary quota into a supposedly universal specification.[9] Check the current dashboard before committing a production workload, and use the $20 launch credit for a constrained evaluation first.

Muse Spark 1.1 vs Claude and GPT: what does the benchmark snapshot say?

In Artificial Analysis' July 12 snapshot, Muse Spark 1.1 scored 51, versus 60 for Claude Fable 5 and 59 for GPT-5.6 Sol.[2] That does not settle every workload. The configurations, reasoning effort, task mix, cost assumptions, and safety policies differ. It does establish the commercial question: can Meta's lower price and one-million-token context overcome a capability gap on the specific agent task a buyer actually has?

The Real Contest: Can Meta Own The Work Between Prompt And Action?

Muse Spark 1.1 is not the model that makes every other model irrelevant. Its 51-point third-party score says the opposite: capability is clustering tightly enough that raw intelligence alone is no longer a defensible product category.[3]

That is precisely why Meta's move matters. It has combined competitive capability, first-party serving, a one-million-token context window, low cache pricing, tool calling, and a consumer distribution machine. No single ingredient is exclusive. The integration is the wager.

The winners will not be the labs that post the largest benchmark number for one week. They will be the platforms that make agents cheap to operate, safe to constrain, easy to evaluate, and native to real work. Muse Spark 1.1 gives Meta a credible seat at that table. The next release has to prove that Meta can own the workflow, not just enter the chart.

What To Do With Muse Spark 1.1

1

Treat the 51 Intelligence Index score as a reason to test, not a reason to standardize. The result is strong but not category-leading.

2

Benchmark completed-task cost, not only token price. The low rates and cache discount are compelling, while reasoning output can be verbose.

3

Use the million-token window for evidence-rich, bounded workflows. Do not use it to bypass retrieval hygiene or prompt-injection defenses.

4

Keep agent permissions narrow. Tool allowlists and workspace isolation should be the baseline for production deployment.

5

Watch whether Meta turns API usage into distribution inside consumer products. That integration, not one benchmark, is the strategic prize.

Sources & References

Primary Meta materials and Artificial Analysis measurements used in this analysis. Vendor claims and third-party figures are separated in the article.

#SourceOutletDateKey Takeaway
1
Meta AI
Meta Superintelligence Labs
Jul. 9, 2026Primary launch announcement: public-preview API, Thinking-mode availability, and product claims.
2
Artificial Analysis
Jul. 12, 2026Configuration-specific Intelligence Index, pricing, speed, latency, modality, context, and token-use data.
3
Artificial Analysis
Jul. 10, 2026Launch comparison, task-cost estimate, benchmark changes, abstention caveat, and pre-release-evaluation disclosure.
4
Artificial Analysis
Accessed Jul. 12, 2026Methodology for the nine-evaluation composite index.
5
CNBC
Jul. 9, 2026Independent reporting on the public preview, initial access, pricing, and $20 credits.
6
Meta Superintelligence Labs
Jul. 9, 2026Vendor safety, robustness, capability results, and recommended deployment controls.
7
Meta Developers
Jul. 9, 2026Developer guidance for OpenAI compatibility, tool use, web search, and model access.
8
TechCrunch
Lucas Ropek
Jul. 9, 2026Independent market framing of Meta's paid coding-agent entry.
9
Meta Developers
Accessed Jul. 13, 2026Authenticated developer documentation for account-specific pricing and rate-limit details.
9 sourcesOpen a linked source to visit the original

Last updated: July 13, 2026