Google announced Gemini 4 Argon on Sept. 30, and the first people outside the company to use it are security teams1. Argon is rolling out to "a set of trusted cyber defenders" through Google's Fairwind Program, wrote Koray Kavukcuoglu, who took over day-to-day leadership of Google DeepMind in August12. Those defenders get a different build from the one the public will see1. "For trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities," Kavukcuoglu wrote1.
Developers, businesses and consumers come later, starting with paid API customers and Google AI Ultra subscribers, on a date Google has yet to set1. The price is already public. Argon starts at $2 per million input tokens and $10 per million output tokens1. It then rises to $4 and $20 once an introductory period ends, on a date Google's post leaves open3.
Google also published a chart comparing Argon across 18 tests with rival models, including OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.53. By VentureBeat's count, Argon leads 12 of them outright and ties one3. Minutes before the announcement, Bloomberg reported that some Google employees find the model weaker in daily work than on benchmarks, especially on certain coding tasks2. Which account holds up? The first outside scores, posted within hours, put Argon level with GPT-6 Astra on one index and first on two others456.
Defenders first
Fairwind opened in September with a smaller model, Gemini 3.8 Flash Cyber7. It has since signed up more than 650 organizations, including CrowdStrike and Palo Alto Networks. Google is also taking part in the U.S. government's voluntary process for reviewing models before release1. Tulsee Doshi, head of Gemini products at Google DeepMind, explained the order of release to CNBC. "Starting this rollout in this way gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible," Doshi said8.
Wiz, the Google-owned cloud security company, has already used Argon in Scan for Good, its free program for finding exposures in public infrastructure7. Google says the model found a critical flaw exposing personal data in healthcare software used by hospitals worldwide, a flaw that earlier frontier models had missed1. Ars Technica's Ryan Whitwam noted that Google has yet to name the software or describe the flaw9.
Anthropic set the pattern this rollout follows. It has kept Claude Mythos Preview, its most advanced model, with a small number of trusted organizations, Agence France-Presse reported10. Argon arrived a day after Sundar Pichai, Alphabet's chief executive, and five other technology leaders signed a voluntary accord at the White House11. President Donald Trump called it "morally binding." Its four commitments, which The AILately.com Desk covered on Sept. 30, include internal cybersecurity controls and an outside assessor for each company's models11. Pichai left the date for a wider release open on X. "We're going to make it available as soon as we can and as safely as we can," Pichai wrote12.
What Google's chart shows
Harvey's Legal Agent Benchmark shows one of Argon's clearest leads, at 19.6 percent against 5.4 percent for GPT-6 Astra and 3.8 percent for Claude Opus 5.53. DeepSWE v1.1, a test of long software engineering jobs, gives Argon 77.9 percent to Opus 5.5's 74.2 percent and Astra's 74.1 percent3. Zapier's AutomationBench puts Argon at 51.3 percent and Opus 5.5 at 42.5 percent3. Security fixes run closer: Argon and Astra tie at 68 percent on CWE-bench v1, one point ahead of Opus 5.53.
Argon trails on the other five, two of them coding tests13. Claude Opus 5.5 leads Terminal-Bench 4.0 at 66.4 percent to Argon's 57.4 percent, and GPT-6 Astra leads FrontierSWE v2 at 65.5 percent to 55.0 percent3. Argon can also write up to 1 million tokens in a single response, up from 64,000 for earlier Gemini models1.
The chart mixes Google's own measurements with numbers its rivals reported. Google computed Argon's DeepSWE score itself, while the rival numbers came from a public leaderboard and company reports, Decrypt's Jose Antonio Lanz reported13. On Terminal-Bench, the chart uses the 66.4 percent Anthropic reported for its own model14. Artificial Analysis measured Opus 5.5 at 59.6 percent and Argon at 57.1 percent on its own setup, Trending Topics reported14.
What outside testers found
Artificial Analysis, an independent benchmarking firm, posted its results 21 minutes after Google's announcement4. Argon scores 53 on its Intelligence Index at the "High" reasoning setting, level with GPT-6 Astra and one point ahead of OpenAI's GPT-6.1 Sol4. Claude Opus 5.5 sits at 58 and Claude Sonnet 5.5 at 56, The Decoder reported6. "Google is now back to being one of the top three labs in intelligence achieved," Artificial Analysis wrote4.
Argon also spends more tokens per task. It used an average of 62,000 output tokens per task, against 27,000 for GPT-6 Astra6. Its hallucination rate on the firm's AA-Omniscience test came in at 15 percent, against 51 percent for Astra6.
Human raters preferred it. Arena, an AI evaluation platform where people compare models in head-to-head votes, put Argon first in its Text Arena with 1,525 points, 20 ahead of Claude Opus 4.65. Its web development board ranks Argon eighth5. Vals AI's index, which weights finance, coding, legal and tax work by each sector's contribution to U.S. GDP, ranks Argon first at 68.9 percent16. It is the first Gemini model to top that index6.
Inside Google
Bloomberg's Julia Love and Davey Alba described a spectrum of opinion inside the company2. Some employees believe Anthropic's and OpenAI's models are improving faster than Gemini, and that Gemini 4 will lag in some areas even at its best. Others believe it has caught up2. People familiar with the internal evaluations called its coding uneven, and one singled out front-end design, the work that shapes how apps and websites look2.
Google disputed the claim that Gemini 4 underperforms in areas such as coding, Bloomberg reported2. A Google employee familiar with model development described a "large consensus" inside the company that the model is at the frontier2. Google also pointed Bloomberg to remarks Kavukcuoglu made the week before, at a conference hosted by The Information. "In my mind, it's a certainty that we are always gonna be at the frontier," Kavukcuoglu said2.
Argon arrives after a promised model missed its launch. Google announced Gemini 3.5 Pro at its I/O conference in May and promised it for June, then dropped it, people familiar with the matter told Bloomberg2. The last Pro model before Argon shipped in February2. Bloomberg Intelligence analyst Mandeep Singh estimates that a training run for such a model can cost as much as $400 million2.
The case for
Google engineers have used Argon for daily work from debugging to large codebase migrations, according to the launch post1. Doshi described that testing to Axios on launch day. "We've seen strong performance up close as Googlers have put the model through its paces in recent weeks, with many relying on it for their hardest coding and research problems," Doshi told Axios15.
Three internal efforts appear in Google's launch post1. A team of Argon agents studied fleetwide profiling data and found memory savings that free more than 300 tebibytes across Google's data centers once rolled out. Other agents are moving C and C++ code to Rust, in projects as large as the Zircon kernel of Google's Fuchsia operating system, at more than 800,000 lines1. For libgav1, Google's open-source video decoder, Argon agents produced a memory-safe Rust version that runs 2.7 times faster than the earlier Rust port, with identical output1. Quantum computing researchers used it to beat a published baseline by 40 percent in minutes, Google says1.
The case against
Bloomberg's reporting points to an industry habit known as "benchmaxxing," in which engineers chase a good test score over a product that does the job well2. Two people familiar with the model told Bloomberg that Gemini 4 appears affected2.
Edwin Chen, founder of the AI startup Surge AI, explained what that habit costs. "An analogy would be, 'Oh yeah, my kid got a really good score on the SAT' — but the SAT doesn't translate into real-world performance," Chen said2. Chen told Bloomberg that relying on benchmarks can lead labs to build models that write code in a particular language2. Apps that are easy to use or well designed lose out, in Chen's account.
Matthias Bastian, who covered the launch for The Decoder, set the first outside scores against Google's chart. "It puts the ad giant back among the top three AI labs, though Anthropic likely still holds the lead," Bastian wrote6.
What it costs
Argon's list price undercuts both rivals for now. At $2 and $10 per million tokens, it charges half of Claude Opus 5.5's rates and a fifth of GPT-6 Astra's $10 and $503. Does the lower token price make Argon cheaper to run? It does for now, by Artificial Analysis's math, and the margin depends on the introductory price4. Heavier token use shrinks the gap. Artificial Analysis puts an average task on its index at $1.99 today, 60 percent of what Astra costs4. Once the introductory price ends, the same task costs $3.98, about 1.2 times Astra's4. One person familiar with its development told Bloomberg that Argon is a very large model, the kind that typically costs more to run2.
What each side needs
Doshi's case needs the coding strength Googlers describe to show up in outside tests once developers get access15. Terminal-Bench 4.0 and FrontierSWE v2 are the place to look, since Google's own chart puts Argon behind on both3. It also needs outside review of the healthcare flaw and the internal results Google has published1. Chen's case needs the public release to reproduce the gap Google employees described to Bloomberg, a model that scores better than it works2. Artificial Analysis gave the first outside reading, with Argon level with Astra and five points behind Opus 5.5, and that reading supports parts of both cases46.
What to watch
- A release date for paid API customers and Google AI Ultra subscribers, the first paying users in line1.
- Outside results on Terminal-Bench 4.0 and FrontierSWE v2, the two coding tests where Google's chart shows Argon behind3.
- An end date for the introductory price; after it, Argon's cost per task moves from 60 percent of GPT-6 Astra's to about 1.2 times4.
- The name of the healthcare software where Wiz found the flaw, and the flaw's details9.
- Arena's Agent Arena ranking, now eighth for Argon on about 3,000 real-world agent sessions and labeled preliminary16.
Sources
- Koray Kavukcuoglu, "Gemini 4 Argon: our next era of frontier intelligence," Google, Sept. 30, 2026, https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
- Julia Love and Davey Alba, "Google Grapples With Employee Skepticism About New Gemini Model," Bloomberg (via Yahoo Finance), Sept. 30, 2026, https://finance.yahoo.com/technology/ai/articles/google-grapples-employee-skepticism-gemini-195242680.html
- Carl Franzen, "Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release," VentureBeat, Sept. 30, 2026, https://venturebeat.com/technology/google-unveils-gemini-4-argon-retaking-benchmark-lead-over-openai-and-anthropic-but-in-limited-release
- Artificial Analysis (@ArtificialAnlys), "Google's new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index," X, Sept. 30, 2026, https://x.com/ArtificialAnlys/status/2105392625788637299
- Arena (@arena), "Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena," X, Sept. 30, 2026, https://x.com/arena/status/2105394855644139908
- Matthias Bastian, "Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead," The Decoder, Sept. 30, 2026, https://the-decoder.com/google-gemini-4-argon-closes-the-gap-with-openai-and-anthropic-but-doesnt-take-a-clear-lead/
- Duncan Riley, "Google's new frontier AI model Gemini 4 Argon goes to cybersecurity defenders first," SiliconANGLE, Sept. 30, 2026, https://siliconangle.com/2026/09/30/googles-new-frontier-ai-model-gemini-4-argon-goes-to-cybersecurity-defenders-first/
- MacKenzie Sigalos and Samantha Subin, "Google rolls out Gemini 4 Argon, its most advanced AI model," CNBC, Sept. 30, 2026, https://www.cnbc.com/2026/09/30/google-gemini-4-argon-ai.html
- Ryan Whitwam, "Google announces Gemini 4 Argon AI model, but you can't use it yet," Ars Technica, Sept. 30, 2026, https://arstechnica.com/google/2026/09/google-announces-gemini-4-argon-ai-model-but-you-cant-use-it-yet/
- Agence France-Presse, "Google restricts access to new AI model over safety concerns," The Straits Times, Sept. 30, 2026, https://www.straitstimes.com/world/united-states/google-announces-gemini-4-flagship-ai-model-after-months-of-delays
- Hadas Gold and Adam Cancryn, "Top AI executives sign commitment to 'self-police' after meeting at White House," CNN, Sept. 29, 2026, https://www.cnn.com/2026/09/29/business/amodei-huang-karp-trump
- Sundar Pichai (@sundarpichai), "Importantly Argon has frontier safeguards and we are rolling it out responsibly," X, Sept. 30, 2026, https://x.com/sundarpichai/status/2105387954474746264
- Jose Antonio Lanz, "Gemini 4 Is Here, and Google's Flagship Tops All Other AI Models on Cybersecurity," Decrypt, Sept. 30, 2026, https://decrypt.co/379784/gemini-4-google-flagship-tops-ai-models-cybersecurity
- Jakob Steinschaden, "Gemini 4 Matches GPT-6 Astra but Trails Opus 5.5," Trending Topics, Sept. 30, 2026, https://www.trendingtopics.eu/gemini-4-artificial-analysis-en/
- Madison Mills, "Google unveils long-awaited Gemini 4," Axios, Sept. 30, 2026, https://www.axios.com/2026/09/30/google-gemini-4
- Arena (@arena), "Gemini 4 Argon (High) is #8 in Agent Arena," X, Sept. 30, 2026, https://x.com/arena/status/2105411271525052418
