# MiniMax M2.5: Frontier AI at 20x Less Than Claude Opus

**Plutonous** | February 12, 2026 | 7 min read

> MiniMax M2.5 - 230B params, 10B active - scores within 0.6 points of Claude Opus 4.6 on SWE-bench at 20x lower cost, backed by a Hong Kong IPO.

Tags: MiniMax, M2.5, Open-Weight AI, Hailuo AI, Chinese AI, AI Pricing, SWE-bench, Agentic AI

---

**TL;DR:** MiniMax released M2.5 on February 12, 2026, an open-weight coding and agentic model that scores 80.2% on SWE-bench Verified (within 0.6 points of Claude Opus 4.6) while charging just $0.15 per million input tokens, 33x cheaper than Opus<sup><a href="#source-1">[1]</a></sup>. The company IPO'd in Hong Kong a month ago at an $11.5B valuation, shares have since quadrupled, and 30% of all tasks at MiniMax HQ are now completed by their own model<sup><a href="#source-2">[2]</a></sup>. This is the first open-weight model to genuinely match Claude Sonnet-tier performance, and it rewrites the economics of AI development.

The Chinese AI labs have been releasing models at a pace that makes Western product cycles look leisurely, but MiniMax M2.5 is different. It's not incrementally better. It represents a structural break in what open-weight models can achieve, particularly for the agentic coding workflows that are driving the largest share of enterprise AI spend in 2026.

While ByteDance was grabbing headlines with Seedance 2.0 video clips and DeepSeek was teasing V4, MiniMax quietly published benchmark results that made the entire open-source community stop and recalibrate. An open-weight model matching the coding performance of the most expensive frontier models at one-twentieth the cost isn't a minor optimization. It's the kind of shift that forces enterprise procurement teams to rewrite their AI budgets.


### Why This Matters Now

MiniMax M2.5 dropped during the Chinese Spring Festival AI blitz of February 2026, exactly one year after the DeepSeek R1 shock. But while DeepSeek proved Chinese labs could match Western reasoning capabilities, M2.5 proves they can match Western *agentic coding* capabilities, the single highest-value commercial AI use case, at a fraction of the price<sup><a href="#source-1">[1]</a></sup>. The OpenHands evaluation team ranked it the #4 model overall, the first open-weight model to ever exceed Claude Sonnet on their composite benchmark<sup><a href="#source-3">[3]</a></sup>.


## The Numbers That Matter: M2.5 By the Benchmarks

Let's be clear about what MiniMax achieved. This isn't a model that trades well on cherry-picked evaluations. The SWE-bench Verified score of 80.2% puts M2.5 within striking distance of Claude Opus 4.6 (80.8%), a model that costs $5.00 per million input tokens versus M2.5's $0.15<sup><a href="#source-1">[1]</a></sup><sup><a href="#source-4">[4]</a></sup>.


- label: SWE-bench Verified; value: 80.2%; description: Within 0.6 pts of Claude Opus 4.6
- label: Multi-SWE-Bench; value: 51.3%; description: First place globally (SOTA)
- label: BrowseComp; value: 76.3%; description: Industry-leading web search
- label: BFCL Tool Calling; value: 76.8%; description: Beats Claude 4.6 and Gemini 3 Pro
- label: Input Price; value: $0.15/M; description: 33x cheaper than Opus 4.6
- label: Active Parameters; value: 10B; description: Of 230B total (MoE)


What's often overlooked is the Multi-SWE-Bench result. At 51.3%, M2.5 holds the #1 position globally on the multi-language coding benchmark, not just among open-weight models but among all models period<sup><a href="#source-1">[1]</a></sup>. The tool-calling score of 76.8% on BFCL outperforms Claude Opus 4.6, Claude Sonnet 4.5, and Gemini 3 Pro. For agentic workflows that depend on reliable function calling, this isn't a marginal difference.


20x

Cost reduction vs. Claude Opus 4.6 per task

M2.5 costs approximately $0.15 per task compared to $3.00 for Claude Opus 4.6, according to ThursdAI independent analysis


## Architecture: 230 Billion Parameters, 10 Billion Active

The uncomfortable truth about why M2.5 is so cheap is also the reason it's so good. MiniMax built on a Mixture-of-Experts architecture with 230 billion total parameters but only 10 billion active per inference pass<sup><a href="#source-5">[5]</a></sup>. This sparse activation means you get the knowledge capacity of a massive model with the compute costs of a much smaller one.


### M2.5 Technical Architecture
- title: Sparse MoE; description: 230B total parameters with only 10B active per token, enabling frontier performance at commodity hardware costs; examples: - 23:1 sparsity ratio
- Efficient self-hosting
- Modified MIT License
- title: Forge RL Training; description: Trained with reinforcement learning across 200,000+ real-world environments including codebases, browsers, and office apps; examples: - CISPO algorithm
- Two-month training period
- Real-world task optimization
- title: M2.5-Lightning; description: Speed-optimized variant that doubles throughput at double the price, still 16x cheaper than Opus; examples: - ~100 tokens/second
- $0.30 input / $2.40 output
- Optimized for latency-sensitive tasks
- title: 200K Context Window; description: Native 200K token context, practical for most coding and agentic workflows; examples: - Full codebase analysis
- Multi-file refactoring
- Long document processing


The model was trained using MiniMax's proprietary CISPO algorithm (Clipping Importance Sampling Policy Optimization), first introduced in their M1 paper<sup><a href="#source-5">[5]</a></sup>. What makes M2.5's training unique is the Forge Reinforcement Learning framework: rather than training on synthetic benchmarks, MiniMax trained across 200,000+ real-world environments, actual codebases, web browsers, and office applications<sup><a href="#source-1">[1]</a></sup>.

Here's the genius of this approach. Traditional benchmark training optimizes for benchmark performance. Forge RL optimizes for the messy, unpredictable environments where AI agents actually need to work. That's why M2.5's BrowseComp score (76.3%) is so strong: the model was literally trained to navigate real websites, not simulated ones.

## The Pricing Bloodbath: Intelligence Too Cheap to Meter

MiniMax isn't being subtle about the economic argument. They're calling it "intelligence too cheap to meter," a deliberate echo of the early nuclear energy promise<sup><a href="#source-6">[6]</a></sup>.


- Input $/1M
- Output $/1M
- SWE-bench

- feature: MiniMax M2.5; values: - $0.15
- $1.20
- 80.2%
- feature: MiniMax M2.5-Lightning; values: - $0.30
- $2.40
- 80.2%
- feature: DeepSeek V3.2; values: - $0.28
- $0.42
- ~75%
- feature: Claude Sonnet 4.5; values: - $3.00
- $15.00
- ~72%
- feature: Claude Opus 4.6; values: - $5.00
- $25.00
- 80.8%
- feature: GPT-5.2; values: - $1.75
- $14.00
- ~79%
- feature: GPT-5.2 Pro; values: - $21.00
- $168.00
- ~82%

0


The math is devastating for Western labs' pricing models. Running four M2.5 agents continuously for an entire year costs approximately $10,000<sup><a href="#source-1">[1]</a></sup>. One hour of continuous M2.5-Lightning operation costs roughly $1. For startups and enterprises building agentic AI products, this isn't a price difference. It's the difference between "we can build this" and "we can't afford to build this."


OpenHands Independent Evaluation, February 2026
A two-horse race between Claude Opus on the most capable but pricy side, and M2.5 on the very inexpensive and still highly capable side.


## The Hallucination Problem: What the Self-Reported Benchmarks Don't Tell You

Here's the contrarian take that most coverage of M2.5 is ignoring. Artificial Analysis, the independent AI evaluation firm, ran M2.5 through their AA-Omniscience benchmark and found an 88% hallucination rate, up from M2.1's already concerning 67%<sup><a href="#source-7">[7]</a></sup>.

Their Intelligence Index places M2.5 at a score of 42, tied with GLM-4.7 and DeepSeek V3.2 for the #3-5 spots among open-weight models. That's behind Zhipu's GLM-5 (50) and Moonshot's Kimi K2.5 (47)<sup><a href="#source-7">[7]</a></sup>.

What this means: M2.5 is genuinely excellent at structured tasks like coding, tool calling, and agentic workflows. But for open-ended knowledge tasks requiring factual accuracy, the model hallucinates significantly more than competitors. This is the classic RL-for-coding tradeoff: heavy reinforcement learning on coding tasks can degrade general knowledge reliability.


### The Hallucination Caveat

MiniMax's self-reported benchmarks emphasize coding and agentic capabilities where M2.5 genuinely excels. But independent evaluation by Artificial Analysis shows an 88% hallucination rate on their omniscience benchmark<sup><a href="#source-7">[7]</a></sup>. For enterprise deployments requiring factual accuracy (legal analysis, medical information, financial reporting), this gap matters enormously. M2.5 is a coding and agentic powerhouse. It is not a general-purpose knowledge oracle.


## MiniMax: From SenseTime Veterans to $11.5 Billion IPO

The company behind M2.5 has one of the most remarkable trajectories in Chinese tech. Founded in December 2021 by Yan Junjie, former VP of SenseTime, MiniMax has moved from stealth to public company in just four years<sup><a href="#source-8">[8]</a></sup>.


### MiniMax Company Timeline
From SenseTime veterans to $11.5B public company in four years

- year: Dec 2021; milestone: Company Founded; innovation: Yan Junjie (ex-SenseTime VP) launches MiniMax in Shanghai
- year: 2022-23; milestone: MiHoYo Backing; innovation: Early investment from the Genshin Impact developer
- year: Mar 2024; milestone: $600M Round; innovation: Led by Alibaba at $2.5B valuation. Tencent, Hillhouse also invest
- year: Sep 2024; milestone: Hailuo AI Launch; innovation: Video generation goes viral globally, puts MiniMax on the map
- year: Sep 2025; milestone: Hollywood Lawsuit; innovation: Disney, Universal, and Warner Bros. file copyright suit in U.S. federal court
- year: Oct 2025; milestone: M2 Model Launch; innovation: 230B total / 10B active parameters. MoE architecture at commodity pricing
- year: Jan 2026; milestone: Hong Kong IPO; innovation: Raises HK$4.8B ($620M), shares surge 110% on debut to $11.5B valuation
- year: Feb 2026; milestone: M2.5 Released; innovation: 80.2% SWE-bench, $0.15/M tokens. Shares jump 15.7% to HK$680


The investor list reads like a who's who of Chinese tech: Alibaba, Tencent, Hillhouse Investment, HongShan (formerly Sequoia China), and IDG Capital<sup><a href="#source-8">[8]</a></sup>. Revenue hit $53.4 million in the nine months ending September 2025, up 174% year-over-year, though the company posted a $512 million net loss over the same period<sup><a href="#source-9">[9]</a></sup>.


- label: IPO Valuation; value: $11.5B; description: Hong Kong Stock Exchange, Jan 2026
- label: Revenue Growth; value: 174%; description: YoY, nine months ending Sep 2025
- label: Global Users; value: 200M+; description: Across Hailuo AI, Talkie, and platform
- label: Net Loss; value: $512M; description: Nine months ending Sep 2025


## The Consumer Empire: Hailuo AI and Talkie

What separates MiniMax from many Chinese AI startups is the consumer distribution. With 200+ million cumulative users across 200+ countries, MiniMax has built a consumer flywheel that most AI labs can only dream of<sup><a href="#source-10">[10]</a></sup>.


### MiniMax Product Portfolio
- title: Hailuo AI; description: Consumer multimodal platform with video generation (Hailuo 02/2.3), text, and music creation; examples: - Text-to-video generation
- Viral social media content
- Global availability
- title: Talkie; description: AI companion and chatbot app with character personalities and entertainment focus; examples: - AI character conversations
- Celebrity personalities
- 200M+ users across products
- title: Speech-02; description: Text-to-speech model supporting 30+ languages with exceptionally long input processing; examples: - Multilingual voice synthesis
- Long-form narration
- Character voice acting
- title: MiniMax Open Platform; description: Enterprise and developer API at platform.minimax.io with pay-as-you-go pricing; examples: - M2.5 and M2.5-Lightning APIs
- Video and speech APIs
- Modified MIT License


The video generation side of the business is what initially put MiniMax on the global radar, but it also brought legal trouble. In September 2025, Disney, Universal, and Warner Bros. (plus Marvel, Lucasfilm, DC Comics, and others) filed a copyright lawsuit in U.S. federal court alleging that Hailuo AI generates copyrighted characters on demand<sup><a href="#source-11">[11]</a></sup>. The lawsuit is ongoing.

## The Spring Festival AI War: M2.5 in Context

M2.5 didn't launch in a vacuum. It dropped during what Chinese tech media is calling the "Spring Festival AI War," a concentrated burst of model releases that coincided with Lunar New Year 2026<sup><a href="#source-12">[12]</a></sup>.


### Chinese AI Lab Landscape (February 2026)
- audience: DeepSeek; impact: V4 imminent. Started the price war with R1 in January 2025. Now at $0.28/$0.42 per 1M tokens. Expanded to 1M token context.; details: - Price war originator
- $0.28/$0.42 per 1M tokens
- 1M token context window
- audience: Zhipu (Z.ai) / GLM-5; impact: GLM-5 leads Artificial Analysis Intelligence Index at score 50. IPO'd alongside MiniMax. Strongest on general intelligence benchmarks.; details: - Intelligence Index #1 (score 50)
- Hong Kong IPO
- General intelligence leader
- audience: Alibaba / Qwen; impact: Qwen 3.5 in preparation. Spending CNY 3B ($434M) on user acquisition. Qwen overtook Meta's Llama in cumulative downloads.; details: - $434M user acquisition spend
- Overtook Llama in downloads
- Qwen 3.5 in development
- audience: ByteDance / Seed2.0; impact: Released Seedance 2.0 for video, full-stack AI ecosystem across LLMs, vision, and video at aggressive pricing.; details: - Full-stack AI ecosystem
- $0.47/M input tokens for Pro
- Cinema-grade video generation
- audience: Moonshot / Kimi K2.5; impact: K2.5 ranks #2 among open weights on Intelligence Index (score 47). Strong math with 96.1% on AIME 2025.; details: - Intelligence Index #2 (score 47)
- 96.1% on AIME 2025
- Math reasoning specialist


What makes MiniMax's position unique in this crowded field is the cost-performance niche. GLM-5 leads on raw general intelligence. Kimi K2.5 leads on math. But M2.5 leads on the metric that matters most to enterprise buyers: coding and agentic performance per dollar spent<sup><a href="#source-3">[3]</a></sup>.

## What This Means: The Open-Weight Tipping Point

The uncomfortable truth for Western AI labs is that M2.5 represents a tipping point for open-weight models. When an open-weight model can match 99.3% of the top proprietary model's coding performance at 3% of the cost, the value proposition of closed-source APIs becomes much harder to justify for pure coding and agentic workloads.


### Key Takeaways for AI Teams
- title: M2.5 is a coding and agentic specialist, not a general-purpose model; description: The 88% hallucination rate on Artificial Analysis benchmarks means M2.5 should not replace general-purpose models for knowledge-intensive tasks. Use it for what it's best at: code generation, tool calling, and agentic workflows.; tip: Run M2.5 for coding and tool-calling tasks, keep a general-purpose model for knowledge-intensive queries
- title: The price point enables entirely new architectures; description: At $0.15/M input tokens, you can run multi-agent systems with dozens of M2.5 instances for less than the cost of a single Claude Opus call. This changes what's architecturally possible.; tip: Prototype multi-agent systems on M2.5 first, then evaluate whether you need a more expensive model
- title: Self-hosting is genuinely viable; description: With only 10B active parameters on a 230B MoE architecture and a modified MIT license, organizations can run M2.5 on-premises at a fraction of the cost of API access to Western frontier models.; tip: Check the modified MIT license requirements: commercial use requires 'MiniMax M2.5' attribution
- title: Watch the copyright litigation; description: The Disney/Universal/Warner lawsuit against MiniMax could set precedent for all AI-generated content. Enterprise users should monitor this closely before building production workflows on MiniMax products.; tip: Consult legal counsel before deploying MiniMax models for content generation in regulated industries
- title: Validate independently before deploying; description: MiniMax's self-reported benchmarks diverge from independent evaluations. Run your own evaluation suite on your specific use case before committing to M2.5 in production.; tip: Use Artificial Analysis and OpenHands independent benchmarks as your baseline, not MiniMax's self-reported numbers


While competitors were chasing GPT-5 on general intelligence, MiniMax spent two months training M2.5 in 200,000+ real-world environments to be the model that actually does the work. At a price point that makes every other frontier model look like a luxury purchase.


### The Bottom Line

MiniMax M2.5 isn't trying to be the smartest model. It's trying to be the most useful model at the lowest price. And for the agentic coding workflows that dominate enterprise AI spending in 2026, it's succeeding. The question isn't whether open-weight models can match proprietary ones on coding tasks. M2.5 just proved they can. The question is how Western labs respond when their pricing moat evaporates overnight.


## Sources

<a id="source-1"></a>
1. [MiniMax M2.5 Official Announcement](https://www.minimax.io/news/minimax-m25): Official benchmark results, architecture details, and pricing for M2.5 and M2.5-Lightning

<a id="source-2"></a>
2. [VentureBeat: MiniMax M2.5 near state-of-the-art at 1/20th the cost](https://venturebeat.com/technology/minimaxs-new-open-m2-5-and-m2-5-lightning-near-state-of-the-art-while): Analysis of M2.5's competitive positioning and pricing disruption

<a id="source-3"></a>
3. [OpenHands: Open-weight models catch up to Claude Sonnet](https://openhands.dev/blog/minimax-m2-5-open-weights-models-catch-up-to-claude): Independent evaluation ranking M2.5 as #4 overall, first open-weight model to exceed Claude Sonnet

<a id="source-4"></a>
4. [Anthropic: Claude Opus 4.6 Announcement](https://www.anthropic.com/news/claude-opus-4-6): Claude Opus 4.6 benchmark results including 80.8% SWE-bench Verified

<a id="source-5"></a>
5. [MiniMax-M1 Paper on arXiv](https://arxiv.org/abs/2506.13585): Original CISPO algorithm paper and M1 architecture details

<a id="source-6"></a>
6. [The Decoder: Intelligence too cheap to meter](https://the-decoder.com/minimax-m2-5-promises-intelligence-too-cheap-to-meter-as-chinese-labs-squeeze-western-ai-pricing/): Pricing analysis and competitive implications of M2.5's cost structure

<a id="source-7"></a>
7. [Artificial Analysis: M2.5 Everything You Need to Know](https://artificialanalysis.ai/articles/minimax-m2-5-everything-you-need-to-know): Independent evaluation revealing 88% hallucination rate and Intelligence Index placement

<a id="source-8"></a>
8. [CNBC: MiniMax doubles in Hong Kong debut](https://www.cnbc.com/2026/01/09/minimax-hong-kong-ipo-ai-tigers-zhipu.html): IPO coverage and valuation details

<a id="source-9"></a>
9. [TechNode: MiniMax IPO surges 110%](https://technode.com/2026/01/09/mihoyo-backed-ai-firm-minimax-jumps-on-hong-kong-debut-market-value-tops-11-5-billion/): IPO day performance and MiHoYo backing details

<a id="source-10"></a>
10. [MiniMax Wikipedia](https://en.wikipedia.org/wiki/MiniMax_(company)): Company background, product portfolio, and user statistics

<a id="source-11"></a>
11. [Variety: Disney/Warner/NBCU sue MiniMax](https://variety.com/2025/digital/news/disney-warner-bros-discovery-nbcu-lawsuit-minimax-chinese-ai-company-1236520395/): Copyright lawsuit details and allegations

<a id="source-12"></a>
12. [CNBC: China AI Lunar New Year war](https://www.cnbc.com/amp/2026/02/13/china-ai-lunar-new-year-bytedance-baidu-tencent-alibaba.html): Context on the Spring Festival 2026 AI model release wave

<a id="source-13"></a>
13. [MiniMax Platform Pricing](https://platform.minimax.io/docs/guides/pricing): Official API pricing documentation

<a id="source-14"></a>
14. [MIT Technology Review: What's next for Chinese open-source AI](https://www.technologyreview.com/2026/02/12/1132811/whats-next-for-chinese-open-source-ai/): Broader analysis of Chinese AI open-source ecosystem


*Last updated: February 12, 2026*

---

*Source: [LLM Rumors](https://www.llmrumors.com/news/minimax-m25-cheapest-frontier-model)*
