AI Models & Platforms

Google Announces Gemini 4 Argon Frontier Model for Coding, Cyber Defense

mm
Add Unite.AI to your preferred sources on Google

Google announced Gemini 4 Argon on September 30, 2026, describing it as the company’s new frontier model for real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. The model is rolling out first to trusted cyber defenders through Google’s Fairwind Program, at an introductory price of $2 per million input tokens.

In a post on The Keyword, Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google, said safely releasing frontier capabilities at this level requires a phased approach. Google is actively engaged in the U.S. government’s voluntary process for pre-release model access while it gradually expands access, and said it will keep gathering feedback from early testers as it iterates on guardrails.

Argon’s introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input token price. A footnote in the announcement states that after the introductory period expires, the price will be $4 per million input tokens and $20 per million output tokens.

Internal Deployment and Expanded Output Limits

Google said Argon is already powering internal workflows, with thousands of employees highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. The company reported several examples from that deployment. Argon is helping Google’s quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications; in one example, it beat the published baseline by 40% in a matter of minutes. A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing more than 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.

Google said Argon agents are also migrating C/C++ codebases to Rust across the company, scaling from tens of thousands of lines in core libraries such as re2 and libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel. Given the criticality of many of these systems, the large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.

For libgav1, Google’s open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running rounds of profile-guided experiments and studying the compiler’s output. Google said the result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output.

To support longer, more complex use cases, Google expanded the model’s output token limit to 1M tokens, up from the previous 64K tokens, a figure the company describes as industry-leading.

Enterprise Benchmarks and Cybersecurity Defense

Google reports that Argon sets a new state of the art on DeepSWE v1.1 with a score of 77.9%, a benchmark measuring performance on real-world long-horizon software engineering tasks. The company describes Argon as the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP, and reports similarly leading performance on Vals Finance Agent v2 for multi-step financial research and Harvey’s Legal Agent Benchmark for legal research and drafting.

On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Google reports Argon ranks first with a score of 51.3%. On LVBench, which measures long video understanding, the company reports a state-of-the-art score of 91.7%, and said Argon can drive professional chart analysis, identify details from long videos, and take action based on a series of documents.

Google said it trained Argon to autonomously find, validate, and patch critical software vulnerabilities, and it will release the model without cyber guardrails to trusted defenders and its own internal teams. Wiz is already using Argon through its Scan for Good initiative, a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. Google said the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.

On CWE-bench v1, which evaluates a model’s ability to remediate security vulnerabilities, Google reports Argon ties for first place with a top score of 68%. The company said Argon also uncovered exposures across complex codebases spanning 20 programming languages on its internal vulnerability benchmark, and outperformed Gemini 3.8 Flash Cyber on Wiz’s internal black-box penetration testing benchmark, which tests a model’s ability to analyze live web systems without source code.

The Fairwind Program, launched by Google on September 2, 2026, is a limited-access program for governments and trusted partners to use the company’s cyber defense tools. Participating organizations agree to operational standards that include limiting access to employees within internal cybersecurity, incident response, or penetration testing teams and deploying protections such as multi-factor authentication. Google said the program has more than 650 participating partners globally, and its first offering paired Gemini 3.8 Flash Cyber with the CodeMender harness to help defenders find, verify, and fix vulnerabilities.

Frontier Safeguards and Staged Rollout

Google said it is strengthening frontier safeguards in four areas before Argon’s broad availability. For misuse defense, the model is designed to refuse harmful cyber and chemical, biological, radiological, and nuclear requests while preserving legitimate dual-use scientific research under Google’s Frontier Safety Framework. The company said it is improving its techniques to monitor the model’s internal activations to spot misuse, and that the safeguards underwent robustness testing by internal and external red teams using manual and automated attack methods.

Google describes Argon as its most resilient model yet against indirect prompt injection, where malicious instructions or context are used to hijack a model’s behavior, and reports leading robustness on Gray Swan’s Indirect Prompt Injection benchmark through automated red teaming and adversarial training.

The company is also deploying misalignment mitigations that monitor Argon’s chain-of-thought and actions and stop execution when necessary. Google said a similar system monitored its training runs and sent alerts to a dedicated incident response team, with careful precautions against feeding the findings back into training. Finally, Google said it is hardening its sandboxed environments by isolating and sealing them before high-risk training or evaluations begin, in line with its agent control roadmap.

Google said feedback from the initial cohort of cyber defenders and trusted testers will help it strengthen its systems before Argon is released to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers, as soon as possible.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.