Google’s latest Gemini generation is Gemini 4, currently represented by the frontier model Gemini 4 Argon, announced on September 30, 2026. The bottom line is straightforward: Argon looks highly competitive for long-horizon software engineering, enterprise automation, cybersecurity, and very long outputs, but it is not yet a proven across-the-board leader - and most users still cannot access it.
Key Takeaways
- Status: Gemini 4 Argon is rolling out first to vetted cybersecurity partners through Google’s Fairwind Program, not the general public.
- Strength: Google reports leading results on 13 of 19 published benchmark rows, including 77.9% on DeepSWE v1.1.
- Important caveat: Independent testing places Argon at 53 on Artificial Analysis’s Intelligence Index, tied with GPT-6 Astra rather than clearly ahead.
- Context: Google has expanded the maximum output from 64,000 to 1 million tokens, enabling unusually long reasoning and generation.
- Practical verdict: Argon is most compelling for controlled, long-running engineering and cyber workflows; wait for broader access and independent replication before making it a default model.
Gemini 4 Argon is best for: security teams, Google Cloud customers, enterprise developers, and researchers who can participate in the Fairwind rollout and need long-context or agentic workflows.
Other Gemini users are best served by: the latest generally available Gemini model until Argon reaches the Gemini API, AI Studio, Gemini app, and standard enterprise channels.
What Gemini 4 is - and what has actually launched
Gemini 4 is a new model generation from Google DeepMind, but the public announcement currently centers on one frontier model: Gemini 4 Argon. Google describes Argon as a model for complex software engineering, professional knowledge work, cybersecurity defense, and sustained reasoning.
The launch is deliberately staged. Initial access is limited to trusted cybersecurity defenders in Google’s Fairwind Program, where the model is being used to identify and fix software vulnerabilities before wider release. Google has said it intends to expand access to developers, enterprises, and consumers as testing and safeguards progress, but the announcement does not provide a firm general-availability date.
That distinction matters. Gemini 4 Argon has been announced and is being tested in production-like settings, but it should not yet be treated as a normal, generally available API model. Reports on launch day indicated that Argon was not yet available through the Gemini app, Google AI Studio, or the standard Gemini API.
Gemini 4 is currently a restricted frontier-model rollout, not a finished mass-market product.
The headline technical change: one million output tokens
Argon’s most striking published specification is its maximum output length. Google is expanding the output limit from 64,000 tokens in the prior generation to 1 million tokens. This is an output limit, not automatically proof that every request also accepts a one-million-token input context.
A million-token response could support extended code transformations, multi-stage research, detailed security analysis, long-running agent traces, and large generated artifacts. It also creates practical risks: costs can escalate quickly, long answers can become less reliable, and downstream systems may struggle to process or validate enormous outputs.
Long context is useful only when paired with retrieval, state management, checkpoints, and tool-level validation. A model that can emit a million tokens is not necessarily better at deciding which 10,000 tokens matter.
What Google’s benchmark table says
Google’s launch materials report that Argon leads outright on 13 of 19 published benchmark rows, ties one, and loses five when compared with GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5. Those are vendor-reported results, so they should be read as a claim about Google’s evaluation setup rather than a settled independent ranking.
| Evaluation | Gemini 4 Argon | Reported comparison | What it suggests |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | Astra 74.1%; Opus 5.5 74.2% | Strong long-horizon software engineering |
| Vibe Code Bench | 91.9% | All compared models near 89 - 90% | Very strong app-building performance, with a narrow margin |
| FrontierSWE v2 | 55.0% | Astra 65.5%; Opus 5.5 62.3% | Not a universal coding leader |
| Terminal-Bench 4.0 | 57.4% | Opus 5.5 66.4%; Astra 58.2% | Weaker relative command-line performance |
| OSWorld-2.0 | 69.2% | Astra 72.6% | Competitive computer-use capability |
| AutomationBench | 51.3% | Opus 5.5 42.5%; Astra 41.4% | Promising enterprise automation |
The coding results are the clearest example of why a single headline score is insufficient. Argon’s 77.9% on DeepSWE v1.1 is an impressive reported result, but it scores 55.0% on FrontierSWE v2 and 57.4% on Terminal-Bench 4.0, where competitors lead. The reasonable interpretation is that Argon is excellent on some forms of long-horizon software work, not that it dominates every coding environment.
DeepSWE v1.1
Argon’s reported 77.9% on DeepSWE v1.1 is its strongest software-engineering headline. The benchmark evaluates long-horizon work across real repositories and multiple programming languages, making it more relevant than short coding questions.
However, the result remains self-reported in the available launch material. Independent replication, identical scaffolding, tool configurations, and cost-normalized comparisons are necessary before treating the margin over competing systems as definitive.
Agentic coding and terminal work
Argon’s reported 55.0% on FrontierSWE v2 places it behind GPT-6 Astra at 65.5% and Claude Opus 5.5 at 62.3%. On Terminal-Bench 4.0, Argon’s 57.4% trails Opus 5.5 at 66.4% and is slightly below Astra at 58.2%.
These results matter more to experienced engineering teams than polished code-generation demos. Agentic coding depends on planning, repository navigation, shell discipline, test interpretation, recovery from failed commands, and knowing when not to modify a codebase. Argon’s uneven scores suggest that workflow design and harness configuration will strongly influence the outcome.
Independent evidence is more restrained
Artificial Analysis reportedly scores Gemini 4 Argon at 53 on its Intelligence Index at the highest available reasoning setting. That matches GPT-6 Astra and Claude Fable 5.1, while Claude Opus 5.5 is reported at 58.
This does not make Argon weak. A score of 53 places it among the frontier systems. It does mean that Google’s claim of broad benchmark leadership should not be converted into a blanket claim that Argon is the best general-purpose model.
Independent reports also place Argon around 57% on an external Terminal-Bench 4 evaluation and eighth on the Arena Agent Arena leaderboard based on a check of 3,417 sessions on October 1. Leaderboards can change rapidly, and early rankings are affected by limited traffic, model routing, system prompts, and uneven availability, but the results reinforce the same conclusion: Argon is competitive rather than unambiguously dominant.
Some independent reporting describes particularly strong performance on long-context GraphWalks, AutomationBench, Vibe Code Bench, cybersecurity tasks, and the Vals Index. Those results are useful signals, but several are based on early evaluations, vendor-supplied configurations, or limited public methodology. They should be treated as evidence of capability areas, not as a final model ranking.
Pricing: promising numbers, but not yet a normal rate card
Early reports put Argon’s introductory API price at $2 per million input tokens and $10 per million output tokens, with cached input reportedly priced at $0.10 per million tokens. After the introductory period, the reported price is expected to rise to $4 per million input tokens and $20 per million output tokens.
These figures should not yet be treated as a finalized, generally available Google pricing commitment. Argon is not broadly accessible, and a standard public pricing page and complete billing policy were not available in the launch information examined.
| Reported item | Early figure | Reported later figure | Confidence |
|---|---|---|---|
| Input tokens | $2 per million | $4 per million | Early reporting; verify at public launch |
| Output tokens | $10 per million | $20 per million | Early reporting; verify at public launch |
| Cached input | $0.10 per million | Not confirmed | Reported, not yet a settled public rate card |
| Output limit | Up to 1 million tokens | Not applicable | Reported launch specification |
At the introductory rate, a single one-million-token output would cost approximately $10 before retries, tool calls, input tokens, and any platform charges. At the reported later rate, the same output would cost approximately $20. That is affordable for high-value automated engineering work but expensive for careless looping or unbounded agent traces.
Cybersecurity is the first real-world proving ground
Google is positioning Argon as a cybersecurity model as much as a coding model. The Fairwind rollout gives selected defenders early access to use Argon for finding software flaws and producing fixes, while Google participates in a voluntary government process for pre-release model access and safety evaluation.
This strategy makes operational sense. Cybersecurity tasks are difficult to benchmark with a single score, but they provide concrete feedback: vulnerability discovery, exploitability assessment, patch quality, regression risk, and whether an agent can operate safely inside a constrained environment.
The same setting also limits what outsiders can conclude. Results from trusted partners may not transfer to ordinary developers, who will use different tools, repositories, permissions, prompts, and review processes. Security performance must be measured not only by vulnerabilities found, but also by false positives, unsafe fixes, data handling, and reproducibility.
What is still unknown
- General availability: Google has not supplied a firm public launch date for the Gemini app, AI Studio, standard API, or broad enterprise access.
- Model family: The available announcement focuses on Argon and does not establish the complete Gemini 4 lineup, including lower-cost, faster, or multimodal variants.
- Technical report: No complete public technical report was available with detailed training data, architecture, post-training methods, system prompts, evaluation harnesses, or contamination analysis.
- Multimodality: Gemini’s brand is associated with multimodal capabilities, but the available Argon material does not fully specify supported modalities, image and video limits, audio behavior, or production tool interfaces.
- Latency: Publicly reliable tokens-per-second and end-to-end latency figures are not yet established. A model optimized for deep reasoning and million-token outputs should not be assumed to be fast.
- Reliability: Early benchmark scores do not establish hallucination rates, calibration, refusal behavior, or performance under long-running production workloads.
How to evaluate Argon when access opens
Organizations should test Argon against their own repositories and workflows rather than importing Google’s leaderboard ranking directly. The evaluation should measure useful work completed per dollar, not just pass rate.
- Freeze the task set: Use representative tickets covering bug fixes, refactors, tests, migrations, documentation, and incident response.
- Keep scaffolding equal: Match tools, repository snapshots, context retrieval, time limits, retry budgets, and human intervention across models.
- Measure completion quality: Record tests passed, review changes requested, regressions, security findings, and whether the solution can be maintained.
- Track economics: Include input, output, cached tokens, tool calls, failed attempts, and human review time.
- Stress long context: Test whether Argon can locate relevant facts in large repositories without being distracted by irrelevant material.
- Set stop conditions: Require approval before destructive commands, production changes, dependency upgrades, or security-sensitive remediation.
Practical implications for developers and enterprises
For software teams
Argon is worth testing first on repository-scale tasks: tracing a bug across services, implementing a feature with broad test coverage, performing a controlled migration, or analyzing a large dependency graph. Its DeepSWE and long-context results suggest potential value where the work requires sustained planning rather than isolated code completion.
Do not replace an existing coding agent solely because Argon leads one benchmark. Its weaker FrontierSWE and Terminal-Bench results indicate that shell-heavy workflows may favor another model. The correct choice should be determined by the team’s repositories, tool harness, review standards, and failure costs.
For security teams
The Fairwind rollout makes Argon especially relevant to vulnerability discovery and remediation pilots. Use sandboxed environments, least-privilege credentials, deterministic snapshots, and mandatory human review for fixes. Treat generated patches as untrusted until tests and security analysis pass.
For API builders
Wait for the official public rate card and service-level details before committing architecture to Argon. If the reported $2/$10 introductory pricing is confirmed, it could be attractive for high-value automation, but million-token output capacity can create runaway spend without strict budgets and output caps.
For ordinary Gemini users
There is no reason to plan around Argon until Google makes it available through the products you use. The announcement describes a staged frontier rollout, not a consumer feature that can be enabled immediately.
Sources
- R&D World: Google’s overdue Gemini 4 Argon reaches the frontier with mixed results for science
- Emergent: Gemini 4 Argon Benchmarks
- Anadolu Agency: Google unveils Gemini 4 Argon
- Los Angeles Times: Google’s latest launch puts its AI ambitions to the test
- ExplainX: Gemini 4 coding doubts versus Google’s benchmark table
- Implicator: Google staff doubt Gemini 4 Argon coding
- CNET: Gemini 4 Argon availability
- Beam AI: Gemini 4 Argon for AI agents
- Yahoo Finance: Google debuts Gemini 4 Argon
- CNBC: Google rolls out Gemini 4 Argon
- SaaS City: Gemini 4 Argon benchmarks and pricing
- Handy AI: Gemini 4 Argon model drop
- Latent Space: Gemini 4 Argon analysis