# Gemini 4 Argon: The Price Is Public, the Access Is Not

**Plutonous** | October 3, 2026 | 7 min read

> Google's Argon rollout separates frontier capability from practical access. Read the introductory pricing, output-token limit and deployment checks that matter.

Tags: Google, Gemini 4 Argon, AI Pricing, AI Access, Cybersecurity, AI Agents, Enterprise AI, Model Evaluation

---

*Cover: generated editorial illustration of a research engine behind a controlled entrance. It represents access restrictions, not measured model performance.*

**TL;DR: Google announced Gemini 4 Argon on September 30 with a 1 million-token output limit and introductory rates of $2 per million input tokens and $10 per million output tokens.<sup><a href="#source-1">[1]</a></sup> Initial access runs through the restricted Fairwind Program.<sup><a href="#source-2">[2]</a></sup> Buyers should evaluate the permissions, eventual bill and accepted work before treating a model announcement as production capacity.**

The real story isn't another intelligence crown. It is the widening distance between a model existing and an organization being able to use it. Argon makes that distance unusually visible: there is enough information to sketch a business case, while access remains a separate commercial and operational question.

This October 3 analysis examines the September 30 announcement. Google says broader availability will begin with paid API customers and Google AI Ultra subscribers, without specifying a launch date.<sup><a href="#source-1">[1]</a></sup> A roadmap can justify preparing an evaluation. It cannot justify promising a customer that a workflow will be available next week.


### Why This Matters Now

Procurement has to answer three different questions: can the model perform the task, can this organization obtain permission to run it, and does the completed result justify its full cost? Argon makes the second question impossible to ignore.


## The Access Gate: An API Price Is Not an API Entitlement

Fairwind prioritizes governments, critical infrastructure and core technology platforms. Its terms prohibit reselling access and require participating organizations to control and track employee use, including phishing-resistant multifactor authentication.<sup><a href="#source-2">[2]</a></sup>

That changes the immediate opportunity. An eligible security team can prepare an application tied to assets it is authorized to defend. An ordinary software startup should prepare a portable evaluation suite and keep its existing production provider. Buying a consumer subscription in anticipation of access is a different decision from securing a production service.

Here's the genius in the distribution strategy, as a business inference: controlled access can let a supplier learn from demanding customers before exposing the same capabilities to a wider market. The customer gets a possible head start. The supplier gains feedback from workflows where a correct result has clear economic value. Neither party gets a guarantee that the eventual general product will have identical permissions.

Our [Astra access analysis](/news/gpt-6-astra-access-economics-followup) identified a related issue with allowances and capacity. Argon adds eligibility to the equation. Treat those as separate procurement fields, even when the brand name stays constant.

## The Price Ladder: Budget for the Rate After the Introduction

Google's stated post-introduction prices are $4 per million input tokens and $20 per million output tokens. The announcement gives no expiration date for the introductory period.<sup><a href="#source-1">[1]</a></sup>


### Google's Announced Argon Token Rates
US dollars per million tokens; announced rates are not confirmation of account access.

- label: Introductory input; value: $2; description: Google-announced input rate.
- label: Introductory output; value: $10; description: Google-announced output rate.
- label: Later input; value: $4; description: Google-announced rate after the introduction.
- label: Later output; value: $20; description: Google-announced rate after the introduction.

Source: Google, September 30, 2026. These are announced token prices, not measured task costs. Tools, storage, taxes and service-specific charges are excluded.


Consider an illustrative job with 100,000 uncached input tokens and 20,000 billable output tokens. Multiplying those quantities by the announced rates produces **$0.40 initially and $0.80 later**. These are arithmetic scenarios, not measured Argon workloads. Tool calls, retries and human review remain outside the calculation.

A business case that works only at the first rate is a promotion-dependent business case. Keep both columns in the budget from the beginning. Record rejected attempts as costs, too: an agent that produces a plausible patch three times before one passes review has generated three bills and one usable result.

The useful metric is expenditure per accepted task. Token price matters inside that metric, but cannot replace it. A lower bill on work that must be redone is not an efficiency gain.

## The Million-Token Limit: More Output Creates a Review Problem

Google explicitly describes the 1 million figure as an **output** limit, up from 64,000. It is not an announcement of a new input-context size.<sup><a href="#source-1">[1]</a></sup>

The distinction matters because reading more material and continuing a longer reasoning trajectory solve different problems. A buyer seeking larger document ingestion should not infer that requirement is satisfied from this headline. A buyer seeking longer autonomous work should ask what stops the process when it becomes unproductive.

Longer execution can be valuable for repository migrations and investigations that require multiple hypotheses. It can also amplify a mistaken premise. The practical response is to define intermediate deliverables: a reproducible finding, a failing test, a reviewed patch, a validated final state. More room to continue is useful only when there is a way to recognize progress.

Measure abandonment as carefully as completion. If an evaluator quietly discards the longest failed runs, the surviving examples will make both reliability and cost look better than the production experience.

## The Benchmark Boundary: A Score Includes Its Harness

CWE-bench v1 lists Argon with Antigravity at 68% programmatic pass@1 and 62% judge-panel pass@1. The evaluator uses 120 tasks, four rollouts per task, high reasoning and a one-hour limit per rollout.<sup><a href="#source-3">[3]</a></sup> These are evaluator-reported results for a particular system, not a prediction for every repository.

Different graders answer different questions. The gap between those figures is a reason to inspect acceptance criteria, not select whichever score looks best. In a local trial, require both the regression suite and the security review to pass. Report the model identifier, harness, tools, reasoning setting, timeout and retry policy with the result.

There is also a pricing discrepancy worth retaining. CWE-bench values Argon's cached input at $0.20 per million; Google's announced 95% input discount implies $0.10 at the introductory input rate.<sup><a href="#source-1">[1]</a><a href="#source-3">[3]</a></sup> Do not silently substitute one basis for the other when discussing evaluation costs.

Vals supplies another useful boundary: its index combines finance, coding, legal and tax tasks using US economic sector weights.<sup><a href="#source-4">[4]</a></sup> That is a benchmark design choice. A hospital's security backlog or a European software company's migration queue will not automatically resemble that mix. Build a workload-weighted score for the actual buying decision.

## The Operating System: Controls Belong in the Business Case

Google's agent-control roadmap combines monitoring with prevention and response.<sup><a href="#source-5">[5]</a></sup> Its Frontier Safety Framework also describes safety-case reviews around relevant capability thresholds.<sup><a href="#source-6">[6]</a></sup> Those are Google's stated processes, not proof that a customer's deployment inherits the same protection.

The operational lesson is to budget for the surrounding system. A generated vulnerability report needs an owner, evidence, a reproduction path and a remediation decision. Wiz's Scan for Good describes validating findings before privately disclosing them to affected organizations.<sup><a href="#source-7">[7]</a></sup> The handoff is part of the product's value, not clerical work that disappears because a model found the issue.

Data handling deserves the same specificity. Google's documentation says CodeMender can retain encrypted session data during an active scan for up to seven days. It separately describes conditions for achieving zero data retention on managed model calls.<sup><a href="#source-8">[8]</a></sup> A model-level assurance does not describe every agent session or connected tool.

Before comparing vendors, draw the entire path from repository checkout to accepted change. Mark who can read source code, who holds credentials, where intermediate artifacts persist and who can authorize a write. That exercise often reveals the real integration cost earlier than another benchmark run.

## The Buying Decision: Prepare a Trial, Preserve the Exit

The uncomfortable truth is that an impressive restricted model can be strategically important without being your next production dependency. An eligible defender should identify a bounded backlog and request the access it needs. Everyone else can prepare the same evaluation without promising a migration date.

Use representative tasks with known acceptance criteria. Save the starting repository, expected checks and human review time. Include tasks the current system already handles cheaply, because a frontier upgrade needs to earn its place rather than absorb the whole queue by default.


### The Missing Number Is Accepted Work

Neither a published price nor a benchmark score tells you how many reviewable results your account can deliver. Keep access, task acceptance, total spend and retention requirements together in the same decision.


Let's be clear: scarcity can make access feel like a prize. Procurement should resist that psychology. The strongest position is a working alternative, a credible evaluation and a budget that survives the introductory period. The model earns the workload after the evidence arrives.


## Sources & References

<a id="source-1"></a>
1. [Gemini 4 Argon announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)

<a id="source-2"></a>
2. [Fairwind Program](https://deepmind.google/fairwind-program/)

<a id="source-3"></a>
3. [CWE-bench v1 methodology and results](https://cwe-bench.com/)

<a id="source-4"></a>
4. [Vals Index methodology](https://www.vals.ai/benchmarks/vals_index)

<a id="source-5"></a>
5. [Securing the future of AI agents](https://deepmind.google/blog/securing-the-future-of-ai-agents/)

<a id="source-6"></a>
6. [Strengthening the Frontier Safety Framework](https://deepmind.google/blog/strengthening-our-frontier-safety-framework/)

<a id="source-7"></a>
7. [Scan for Good](https://www.wiz.io/scan-for-good)

<a id="source-8"></a>
8. [Agent Platform and zero data retention](https://docs.cloud.google.com/gemini-enterprise-agent-platform/resources/zero-data-retention)


*Last updated: October 3, 2026*

---

*Source: [LLM Rumors](https://www.llmrumors.com/news/gemini-4-argon-access-pricing)*
