Google releases Gemini 4 Argon AI, claiming leadership in network security and software engineering
The Cryptonomist
1h ago
Ai Focus
Google launches new flagship AI model Gemini 4 Argon, claiming it is superior to its predecessors in software engineering, enterprise knowledge management, and network security defense, and will be gradually made available to trusted cybersecurity partners.
Helpful
No.Help

Google has unveiled the Gemini 4 Argon, a new flagship AI model launched by the company. Google claims that this model outperforms its predecessors in software engineering, corporate knowledge work, and cybersecurity defense. The Alphabet system was released on Wednesday, September 30, 2026, but it was not made publicly available immediately; instead, it was first opened to a small group of trusted cybersecurity partners. This move marks the biggest step the company has taken so far towards what it calls "cutting-edge intelligence" and comes at a time when Google is under pressure to prove that its models can compete with those of its competitors in high-risk real-world tasks.

Gemini 4 Argon Leading Performance in Software Engineering and Corporate Work

Google stated that Gemini 4 Argon has demonstrated "cutting-edge" performance in complex real-world workflows involving coding, legal and financial research, as well as security defense. The company said that the model has already been in use within internal engineering teams, and employees have reported improvements in professional coding tasks, deeper research capabilities, and the quality of written work.

Refresh the standard benchmark scores

On the DeepSWE v1.1 benchmarks that measure the performance of long-term, real-world software engineering, Argon achieved a score of 77.9%, which Google claims is the new industry leader. In addition to coding, this model also ranked first in Vals Index. This index tracks the economic impact of finance, coding, law, and tax work, and weights it according to each industry's contribution to the US GDP. CNBC reports that Argon leads OpenAI in this index by 5.1 over GPT-6, Astra, and Anthropic's Fable. The model also ranked first in Zapier's AutomationBench, a test that measures end-to-end execution capability across core business functions, with a score of 51.3%; in LVBench, it achieved 91.7%, which is a benchmark for testing long-video understanding capabilities, and Google says this score is also at the industry lead. Argon's advantage in visual reasoning also extends to professional chart analysis and document-based decision-making, which is very useful for knowledge workers who need to handle large amounts of reports and multi-page documents.

The importance of these results lies in the fact that they demonstrate that this model is not merely a tool dedicated to coding. A system that performs well in legal drafting, financial research, and video understanding indicates that Google is positioning Argon as a solution for corporate clients who need to use the same model across multiple departments, rather than as a collection of discrete, specialized tools.

Cybersecurity defense in collaboration with Wiz

Argon has been specifically trained for enhancing network defense, and Google claims it can autonomously discover, verify, and fix critical software vulnerabilities. In the CWE-bench v1 that assesses the model's ability to repair security flaws, Argon tied for first place with a score of 68%, continuing the performance of Google's 3.8 Flash Cyber model released earlier in September 2026. The latter led in the CWE-bench v0. CNBC reports that Argon also tied with OpenAI's GPT-6 Astra and Grok at 4.7 on the network security assessment benchmark, indicating that it is a direct competitor in this rapidly changing field of AI.

For trusted defenders and Google's internal teams, the company plans to release Argon without any network security safeguards, allowing them to utilize the model's full range of cutting-edge defense capabilities.

Practical applications of independently discovering vulnerabilities

Security company Wiz has adopted its Scan for Good plan that utilizes Argon. This plan is a free initiative aimed at protecting critical public infrastructure by identifying and fixing high-risk vulnerabilities. During an early test, the model detected a serious flaw that could expose sensitive personal information in medical software used by hospitals around the world – Google stated that previous cutting-edge models failed to identify this risk at all.

On Google's internal vulnerability benchmark, Argon identified exposure issues in code repositories covering 20 programming languages; on the Wiz black-box penetration testing benchmark, it also outperformed 3.8 Flash Cyber in terms of identifying attack surfaces and generating proof of concept evidence.

Technological innovation: From output at the million-level token to quantum optimization

To support longer and more complex tasks, Google has significantly increased the output limit of Argon to token to an industry-leading 1 million, up from the previous 64,000. This larger capacity allows the model to generate hundreds of thousands of token in a single inference process, thereby providing deeper processing capabilities for problems that previously required multiple rounds of back-and-forth processing.

This model has already produced measurable results within Google's own infrastructure. In quantum computing research, Argon helped optimize the spatio-temporal resources of several subroutines—namely, the multiplication of quantum bits by gate operators—which are bottlenecks in key applications; in just a few minutes, it improved the public baseline by 40%. Additionally, a team composed of agents from Argon analyzed telemetry data across Google's data center fleet, automatically identified and applied memory optimizations, and after full deployment, over 300 TiB of memory were freed up, with an overall estimate of savings ranging from 500 TiB to 1 PiB. CNBC noted that these improvements did not require the company to purchase additional hardware.

The Argon proxy is still working on one of the most tedious tasks in software engineering: migrating C/C++ code libraries to Rust. This work involves tens of thousands of lines of code in core libraries such as re2 and libgav1, and even extends to over 800,000 lines of code within the Fuchsia OS Zircon kernel. Given the importance of these systems, Google states that the rewritten code undergoes strict automated and manual audits, simulation tests, and reviews before being deployed into production environments. For Google's open-source video decoding library libgav1, the Argon proxy based on an existing Rust porting version replaced 32,000 lines of SIMD code through repeated configuration file-guided experiments, resulting in a decoder that is 2.7 times faster than the previous Rust version while maintaining exactly the same video output.

Frontier security measures before broader opening up

Before Argon is made available to a wider range of users, Google stated that it is strengthening its protective measures. This model is designed to reject harmful requests related to the misuse of networks or CBRN (chemical, biological, radioactive, nuclear), while continuing to support legitimate dual-use scientific research under the company's Frontier Safety Framework. Google claims that it has improved the technology for monitoring the internal activation state of the model in order to detect such misuse, and has tested these protective measures using both manual and automated attack methods by internal and external red teams.

Google also describes Argon as its most resilient model to indirect hint injection attacks to date. Indirect hint injection refers to the practice of hiding instructions in an attempt to hijack the behavior of a model. Through adversarial training and automated red-team testing, the company claims that Argon is ahead of Gray Swan based on the Indirect Prompt Injection benchmark.

This layered approach reflects a broader tension in cutting-edge AI development: the same ability to enable models to discover and fix vulnerabilities can also be misused by unscrupulous individuals to uncover new vulnerabilities. Google's decision to first hand over Argon to a select group of trusted cybersecurity partners, rather than releasing it directly to the public, also indicates the company's awareness of this dual-use risk.

Phased rollout and subsequent arrangements

Google has not fully activated Argon all at once. The company stated that it will release the model in phases, initially targeting a group of network defenders and trusted testers. Their feedback from the real world will help the company continue to refine the system before it reaches developers, enterprises, and consumers.

CNBC also reported that Google was conducting a security assessment with the U.S. government before the release. The timing of the release is also quite significant: just one day before the release, Google's CEO Sundar Pichai signed a voluntary AI security agreement with U.S. President Donald Trump. Prior to this, both parties met at the White House with several technology executives to discuss the growing AI security concerns.

This release also marks a strategic shift. CNBC stated that over the past year, Google has mainly focused on promoting faster and more cost-effective “flash” models, while Gemini and Argon represent the company’s renewed push towards the forefront of innovation – it has been nearly a year since Gemini that the company returned to the forefront of the AI model competition. Whether Argon can maintain this position depends on its actual performance in the hands of external developers and corporate clients, rather than just in Google’s internal testing environments.

This article was generated with the assistance of artificial intelligence and has been reviewed by an editorial team.

Tip
$0
Like
0
Save
0
Views 25
CoinMeta reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
ASTS Notice: AST SpaceMobile Investors can apply to lead a class action lawsuit regarding securities. The deadline is November 13, 2026.
Robbins Geller Rudman and Dowd LLP announce that investors who purchased or acquired AST SpaceMobile securities between March 4, 2025, and July 15, 2026, may apply to serve as the lead plaintiff in a class action lawsuit before November 13, 2026. The complaint alleges that AST SpaceMobile and certain of its executives made false or misleading statements regarding capital needs, liquidity, competitive position, and the adoption rate among users in the United States and Japan.
PR Newswire
·2026-10-01 10:04:44
9
Lloyds and Visa completed weekend stablecoin settlement tests: $750,000 was credited in less than an hour
Lloyds Banking Group of the UK and Visa disclosed a cross-border settlement pilot on September 30: The two parties used USDC to complete an inter-bank obligation settlement worth $750,000, and the funds reached the US end at Visa in less than an hour during the test, which included weekends. For those accustomed to waiting on weekdays and relying on multiple correspondent banks for cross-border payments, the reduction in time is the most noticeable aspect. However, this is a pilot with limited scale involving real funds, and it cannot be claimed that a stablecoin remittance service available to all customers has been fully launched.
币界网
·2026-10-01 09:58:12
15
34 million businesses in the EU support the economy: Large companies account for only 0.2%, yet they generate nearly half of the added value
Eurostat released the final figures for 2024 structural business statistics on September 30: The EU's business economy consists of approximately 34 million enterprises, employing 164.5 million people including self-employed individuals, with turnover exceeding 38.7 trillion euros and generating a value added of 10.9 trillion euros. What is more noteworthy is not the number of enterprises themselves, but rather their size distribution: Large enterprises with over 249 employees account for only about 0.2% of the total number of enterprises, yet they contribute about 49% of the value added and 37% of employment. The scale effect remains a key to understanding Europe's productivity and competitive landscape.
币百科
·2026-10-01 09:58:09
9
UK Q2 GDP Revised Up to 0.5% Growth: Consumption Improves, but Recovery Remains Uneven
On September 30, the Office for National Statistics (ONS) in the UK released the final figures for the national accounts for the second quarter, revising the quarter-on-quarter growth rate of real GDP from the preliminary figure of 0.4% to 0.5%. This change indicates that the economic performance in the spring was slightly better than initially estimated, but it does not alter the overall picture of modest growth: the growth rate was 0.6% in the first quarter and slowed slightly in the second quarter; while the services and construction sectors expanded, the production sector contracted slightly. For those observing the UK economy, the sectoral differentiation behind these revised figures is more noteworthy than the additional 0.1 percentage point.
币百科
·2026-10-01 09:58:07
10
NVIDIA Launches Agent Security Platform: To Ensure That AI Acts Within Defined Boundaries Before Taking Action
As AI intelligents move from answering questions to operating files, invoking tools, and accessing enterprise systems, security issues have also shifted from "saying the wrong thing" to "doing the wrong thing." On September 28th, NVIDIA announced Open Agent Safety Platform, aiming to isolate, monitor, and handle the operations of these intelligents within the same open architecture. The announcement includes an open-source runtime OpenShell, as well as a reference design for Sentry geared towards independent monitoring. The focus of the product is not to ensure that every answer given by the model is correct, but rather to limit the actions that the model can take even if it is misled.
CoinMeta
·2026-10-01 09:58:05
11
View More