
Claude Code Weekly Limits After September 14: What Actually Changed
Anthropic ended a 50% Claude Code weekly-limit promotion and kept a permanent 25% uplift. The practical change is smaller than a '25% cut' headline suggests.

House analyst byline
Plutonous is the LLM Rumors house byline for research-driven analysis shaped by the editorial desk. It keeps the site accountable to one consistent voice while letting individual tips, edits, and research contributions stay behind the piece.
88 published stories

Anthropic ended a 50% Claude Code weekly-limit promotion and kept a permanent 25% uplift. The practical change is smaller than a '25% cut' headline suggests.

A worked guide to OpenRouter Batch API costs, deadlines, ambiguous submissions, result reconciliation and data retention.

A precise check for Apple Intelligence, Siri AI, language, device, platform and EU availability before you expect the new assistant.

Separate on-device AI, Private Cloud Compute and ChatGPT. Check what works offline, inspect Apple's request report and understand iOS 27 availability.

Calculate Claude prompt caching's break-even point, choose a TTL from actual reuse, and diagnose misses without confusing API savings with subscription limits.

DeepSeek V4.1 Flash changed aliases, vision support and API behavior. Use this migration checklist to test cache boundaries, tool reasoning, output limits and user isolation.

METR's time horizon measures human task difficulty, not an agent's unattended runtime. Here is how to read the uncertainty and design a business acceptance test.

What OpenRouter's 50 and 1,000 daily free-request limits actually buy, why paid and free privacy settings differ, and how to avoid surprise fallback costs.

The most-read Astra stories are not only about intelligence. They are about shared allowances, queue design and who gets dependable access to a frontier model.

PrismML packages Bonsai 2 27B as a 5.95 GB GGUF language model or an 8.60 GB MLX multimodal release. The deeper story is how ternary weights change local AI economics.

TypeSafe AI's Jev model returns typed decisions and calibrated probabilities instead of prose. Here is where that design could change software automation, and where the claims remain unproven.

Union Alpha has been revealed as Circuit & Chisel's Pareto 26.9. Its architecture, pricing and uneven benchmark profile show why the product is the orchestration layer, not a secret foundation model.

Claude Opus update rumors now include 5.2 as well as 5.1. What is sourced, what remains speculation, and what a meaningful everyday rival to Astra would need.

Reduce avoidable Astra usage in Codex by checking Fast mode, matching reasoning to the task, controlling context and delegating deliberately. Keep plan limits, credits and API bills separate.

GPT-6 Sol launch rumors point to OpenAI's next model question: can Astra-level progress become a practical everyday tool? Dates, pricing and performance remain unconfirmed.

Apple rates US eSIM-only iPhone 18 Pro models at up to 36 and 45 hours of video playback. Germany lists 34 and 43. Here is what those controlled ratings do and do not tell a buyer.

Union Alpha is a free anonymous OpenRouter model with image input and tool support. Its listing is real; links to GPT-6 Sol or a new Claude Opus remain unproven.

ChatGPT for Financial Services pairs GPT-6 Astra with licensed data, citations, templates, and controls. The product’s real test is whether every conclusion remains reviewable.

A close reading of three new arXiv papers separates their mathematical claims, their authors' AI-use statements, and the evidence readers need before calling AI-assisted mathematics established.

Amodei's new pacing proposal turns AI safety from a lab promise into an argument about who gets to inspect the lab. The split between embedded evaluators, competitor review, and regulatory coordination is the real fight.

METR has disclosed meaningful safeguards, but its public record leaves key questions about current governance, non-cash dependence and publication rights unanswered.

A source-led examination of the people who help shape METR's evaluations, its historical ARC links, and the distinct intellectual networks often collapsed into one story.

DeepSeek-V4.1-Flash pairs 1M-token context with a $0.006-per-million peak cache-hit price. Explore the architecture, API changes and economics of reusable agent context.

What Habitat actually does, explained through a railway station: fetching application data, directing requests and keeping slow dependencies from holding everything up.

Compare A20 Pro with A19 Pro, Snapdragon 8 Elite Gen 5 and Tensor G6: CPU, GPU, Neural Engine, memory bandwidth and cooling, with vendor claims labeled.

A practical guide to GPT-6 Astra access on ChatGPT Plus, including Work and Codex allowances, reset windows, effort settings, Fast mode, and troubleshooting.

Explore an interactive 3D phone chip and learn how CPU, GPU, Neural Engine, memory, quantization and cooling work together to run AI on your device.

iPhone Duo starts at $1,999. Explore US release dates, AirPods 5 pricing, A20 Pro and what Desert Ant and Liquid AI mean for on-device apps.

What Navier-Stokes means, what OpenAI's forced blowup proof claims, and why the dispute over credit and research data matters for AI science.

Anthropic's J-space research gives auditors a new way to inspect silent intermediate reasoning. Mixed replication results show why an internal readout still needs a behavioral test.

Show Astra a picture of sound, and it can offer a surprisingly specific guess. Here is how spectrograms make that possible, and what it would take to turn the demo into a useful tool.

E2B, Daytona, Modal and Cloudflare expose different controls for agent execution. The harder boundary is what a contained agent is authorized to do.

How AI labs collect web data, generate synthetic examples, and turn agent actions into training signals. The bottleneck is verifying what those examples teach.

GitHub HydraFusion, Sakana Fugu, and OpenRouter Fusion point to a market where the durable advantage is the runtime policy that chooses, escalates, critiques, and synthesizes models.

From editable steam trains to a playable portfolio and an iPod-style Mac app, Astra's early builders are turning personal ideas into software worth exploring.

OpenAI's GPT-6 Astra can operate software, carry long-running context, and tackle harder cyber work. The real contest is now the system around the model.

Grok Bot turns AI from a chat window into a persistent operator. That makes its convenience real, its backlash rational, and authority design the next product battleground.

Wafer, RunInfra, OpenRouter's provider price war, and AI-written GPU kernels are turning inference optimization into the next strategic control plane.

Tencent's Hy4 preview pairs a 770B-parameter sparse backbone with Apache-2.0 weights. Together with Z.ai's GLM releases, it shows why open-weight AI is becoming a strategic market, not a side project.

Z.ai's GLM-5.3 and GLM-5.3-Flash show how Chinese labs can compound public weights, DeepSeek and Moonshot research, shared RL stacks, and lower-cost deployment.

OpenAI's postmortem describes 1,200 agents, 70,000 messages, and a real Hugging Face intrusion. The lesson is about containment, incentives, and agent operations.

Z.ai revealed that Ox Alpha was an early GLM-5.3-Flash preview. The anonymous experiment drew 343.5 million OpenRouter requests before the company attached its name.

OpenAI says Jalapeño delivers up to 1.9x more work per watt and 3.6x lower latency than selected Blackwell systems. The real threat to NVIDIA is inference share and pricing power, not a 2026 GPU collapse.

Ox Alpha reached 70.3 million OpenRouter requests and 5.93 trillion prompt-plus-completion tokens in one day. Its maker is still anonymous, and the fine print matters.

SpecPTC launches safe tool calls while agent code is still streaming. Alex Zhang reports 1–1.2x RLM gains, but the bigger story is runtime scheduling.

Grok 4.6 reaches a 61 Intelligence Index score at $2/$6 API pricing. Grok Bot adds persistent computer use, shared state, routines, and a new governance problem.

Prime Intellect's recursive harness framing explains why AI agents are becoming an orchestration, verification, and training-data business, not just a model race.

Alibaba's 2.4T-parameter Qwen3.8-Max pairs 1M context, $2/$6 API pricing, and vendor-reported agent gains with a promised checkpoint few teams could deploy themselves.

Reasoning effort is not a universal quality slider. It is becoming the control plane that allocates tokens, latency, tools, and verification across AI workloads.

TileRT is not just a faster inference engine. Separate Xiaomi MiMo and Z.ai deployments show why ultra-low-latency runtimes are becoming a new battleground for frontier AI products.

Kimi K3 pairs a 57.1 Artificial Analysis score with 2.8T parameters, 1M-token context, $0.94 task cost, and open weights promised for July 27.

Meta Muse Spark 1.1 API pricing, 1M-token context, benchmarks, coding agents, and what Meta's paid agent platform means for developers.

Kolmogorov-Arnold Networks replace scalar weights with learned functions. Two years of evidence show where KANs work, where they fail, and why the idea survived.

Loop Engineering turns the hidden management work around coding agents into software: triggers, scoped execution, independent verification, durable state, budgets, and explicit stop conditions.

Grok 4.5 combines a 54 Intelligence Index score, 90-token-per-second measured speed, $0.31 benchmark task cost, Cursor-trained agent behavior, and live search in xAI's strongest model release yet.

DeepSpec turns speculative decoding from a hidden serving trick into an open training stack, with DSpark claiming 60% to 85% faster V4-Flash generation.

OpenAI's GPT-5.6 Sol, Terra, and Luna launch is not just a model update. It is a preview of AI releases where capability, price, safety, and government access are bundled together.

Huawei's Tau Scaling Law and Intel's 18A-P roadmap show the same semiconductor shift from opposite sides: future chips will be won through systems, not node names alone.

Recursive's automated AI research system is not just a benchmark win. It is a preview of research loops that propose ideas, write code, run experiments, validate results, and keep going.

DeepSWE shows closed labs still lead frontier coding agents, but open-weight models are starting to price the infrastructure layer. That is exactly how Linux won.

DiffusionGemma is not just Google's 4x faster text generation experiment. It is the open-weights counterpunch to Inception's closed Mercury 2 thesis for real-time AI subagents.

SpaceX's $60 billion stock deal for Cursor turns a coding editor into strategic AI infrastructure. The real story is Composer, Colossus, Grok, and the race to own developer work.

The US government ordered Anthropic to suspend Fable 5 and Mythos 5 access for foreign nationals, forcing a global shutdown. The real story is frontier AI becoming controlled infrastructure.

Claude Fable 5 is not just Anthropic's strongest public model. It is the first Mythos-class release built around public access, fallback routing, data retention, and trusted capability gates.

iOS 27 makes Siri AI a private system agent, not just a smarter assistant. Apple's update ties app actions, Camera mode, speed gains, Liquid Glass, and the EU delay into one AI control-plane strategy.

NVIDIA's GTC Taipei at COMPUTEX 2026 was not a normal product keynote. It was a full-stack argument for AI factories, Windows-native agents, Taiwan manufacturing, and physical AI.

Liquid AI's LFM2.5-8B-A1B shows why the small model race is moving from parameter count to active compute, memory bandwidth, local runtimes, and the devices that can run agents at the edge.

World models began as Schmidhuber's curiosity-driven controllers. Genie 3, DreamerV3, GAIA-1, GameNGen, V-JEPA, and open-ended play show why simulation is becoming AI's next platform layer.

Sakana AI and NVIDIA's TwELL shows why sparse LLMs were not blocked by theory. They were blocked by GPU execution economics.

RLMs treat prompts as environments, not inputs. The MIT paper behind Recursive Language Models, the REPL execution loop, and why the AI industry is adopting it.

From a $1B nonprofit pledge to a $300B for-profit empire: the decade-long origin story of OpenAI, from founding through ChatGPT to $12B in annualized revenue.

Qwen3.5-397B-A17B scores 88.4 on GPQA Diamond at $0.60/M, yet was absent from Anthropic's distillation report. That absence explains more than the names that were included.

Anthropic named DeepSeek, Moonshot, and MiniMax for industrial-scale Claude distillation. Evidence is real, but calling it an 'attack' ignores how AI was built.

Anthropic was mocked as the AI company that couldn't ship. Claude 2 was the punchline. Four years later, they own enterprise AI at a $380B valuation.

Google doubled ARC-AGI-2 from 31% to 77% in one update. Gemini 3.1 Pro leads 13 of 16 benchmarks at $2/M tokens, undercutting Claude Opus 4.6 by 60%.

Claude Sonnet 4.6 is the new claude.ai default - preferred over Opus 4.5 59% of the time, with 1M context and $3/$15 per million tokens.

OpenClaw hit 190K+ GitHub stars before OpenAI acquired it - despite 512 security vulnerabilities including an 8.8 CVSS remote code execution flaw.

ByteDance's Seed2.0 reveals a complete AI ecosystem - frontier LLMs, multimodal vision, agentic coding, and cinema-grade video at a fraction of Western pricing.

MiniMax M2.5 - 230B params, 10B active - scores within 0.6 points of Claude Opus 4.6 on SWE-bench at 20x lower cost, backed by a Hong Kong IPO.

Claude Opus 4.6 commands 40% of enterprise AI spend, found 500+ zero-day vulnerabilities, and Claude Code hit $1.1B ARR as Anthropic raised $30B.

xAI's Grok-4 sets new benchmarks in reasoning, coding, and scientific knowledge. Can it maintain its edge as data access costs and energy prices keep rising?

As the 'One Big Beautiful Bill' reshapes US energy policy, UAE and Singapore pioneer integrated AI infrastructure approaches America could adapt.

AI's growth is reshaping global utilities - 66B liters of annual water use, fusion-powered data centers, and companies still paying the carbon bill.

From McCulloch-Pitts' 1943 logical calculus to GPT-4: trace 82 years of neural architecture evolution and how foundational insights led to the transformer.

Tencent's Hunyuan-A13B uses sparse expert activation (80B total, 13B active) to deliver competitive performance at dramatically lower compute cost.

OpenAI's move to Google's TPUs for inference signals a fundamental change in AI compute economics - and why it matters more than Nvidia's stock price suggests.

Sakana AI built an AI that improves itself by editing its own code. Here's how it works and why it matters for the future of AI development.

How Anthropic built a pipeline scanning millions of books for Claude AI training, the legal battles that followed, and what it reveals about AI data infrastructure.