Google launches the Gemini 4 Argon flagship AI model, but its practical performance is questioned by employees, according to reports.
The Block
1h ago
Ai Focus
According to Bloomberg, Google has begun to gradually roll out Gemini 4 Argon, but some internal employees and informed sources have raised doubts about its actual performance on core tasks such as coding. Google, however, asserts that these claims are inaccurate and states that the model is designed for complex tasks in software engineering, finance, law, and cybersecurity.
Helpful
No.Help

IT Home, October 1st news: According to Bloomberg, Google, a subsidiary of Alphabet, has begun to gradually launch the highly anticipated flagship artificial intelligence model Gemini 4. However, there are still doubts within the company regarding the actual performance of this model in core areas such as code writing.

Google made this model available to a small group of selected, trusted cybersecurity partners on Wednesday local time, and stated that it will expand the scope of availability after completing more tests, giving priority to paid subscription users. Google claims that Gemini achieved top scores in several benchmark tests, with one test measuring security capabilities even surpassing that of OpenAI's Astra model.

However, some internal employees stated that these evaluation indicators do not reflect the entire situation. Several individuals directly involved in the project revealed that although the Gemini 4 model performed well in widely used benchmark tests, its actual performance when used by Google employees was not as satisfactory. These individuals, who preferred to remain anonymous and did not wish to discuss internal matters, mentioned that the model encountered difficulties when handling certain coding tasks.

After Google released the above news, the after-hours stock price of Alphabet increased by about 1.7%. However, as Bloomberg reported employees' doubts about the model near the close of regular trading days, the gain seen earlier on Wednesday declined again.

Recently, Google has been working hard on developing models in the hope of competing with OpenAI and Anthropic. People familiar with the matter said that the company originally planned to release another version in June, namely Gemini 3.5 Pro, but later abandoned this development effort.

Google stated that the claim that Gemini 4 performs poorly in areas such as code is not accurate. The company referred Bloomberg's inquiries to Coray Kavukcuoglu, the person in charge of Google DeepMind, who stated last week that he was encouraged by the performance of this model.

"I have full confidence in this team," said Kavukulo at a conference hosted by the technology news website "The Information." "In my opinion, we will undoubtedly continue to be at the forefront of technology, and that is certain."

There are differing opinions within Google on this matter. Some employees believe that the model iteration and improvement speed of Fable and OpenAI of Anthropic has surpassed that of Gemini. In their view, even if Gemini performs at its best, it will still lag behind competitors in certain areas. On the other hand, another group of employees believes that this new version, which is about to be launched, has already caught up with leading artificial intelligence laboratories.

A Google employee familiar with model development work stated that there is already a "widespread consensus" within the company that Gemini 4 is at the forefront of the industry. This individual claimed that Google has conducted rigorous tests on this model and denied that it is unable to handle the complex and messy coding tasks found in reality.

Google urgently needs the success of Gemini. Almost all of Google's products rely on this series of models as their underlying support: this includes not only the artificial intelligence responses provided at the top of Google's main profit-making search pages but also services such as Maps, Gmail, and the browser Chrome. Each of these products has a user base of over one billion, which is a distribution advantage that many competitors do not possess.

However, OpenAI and Anthropic no longer merely sell models; instead, they have turned to developing their own products, which include programming intelligents. If Google cannot produce a top-tier new generation model, their competitors will gain more opportunities to convince ordinary consumers, developers, and various enterprises that future search and software should be built on their opponents' platforms.

In response to questions from Bloomberg, Google stated that although the previous generation of Pro model was released in February, Google's artificial intelligence-related products have continued to grow since then. This includes the enterprise version of Gemini, chatbot applications for general users, as well as the artificial intelligence features within Google Search; the number of users for the latter two products has already exceeded one billion.

Gemini 3

Last November, Google officially launched Gemini 3. This model received positive reviews, and it is generally believed by the industry that this was a turning point in Google's efforts to catch up with OpenAI and Anthropic. At Google's developer conference I/O in May, Google announced the new generation of iterative versions, Gemini 3.5 and Pro, and promised to release them officially the following month. However, the scheduled release date has passed, and sources familiar with the matter say that Google has since put the Gemini 3.5 and Pro projects on hold.

Abandoning this project would not only hinder Google's development plans for artificial intelligence but also likely result in significant time and financial costs for the company. Mandeep Singh, an industry research analyst at Bloomberg, noted that the cost of a single training run for such a large model can reach up to $400 million (Note from IT: Current exchange rate is approximately 2.687 billion RMB); coupled with the labor input of highly paid AI researchers, the overall cost would further increase.

Recently, the development of Gemini 4 has encountered various difficulties one after another. People familiar with internal evaluation work have revealed that the coding skills of this model vary greatly. One informed source stated that Gemini is not adept at front-end design, which is related to determining the appearance and user experience of applications and web pages. This could become a serious setback, as Google is already in a catching-up position in the highly competitive field of artificial intelligence coding tools.

In addition, informed researchers and developers stated that it is a model of very large scale. The operating costs of large models are usually high, which may put pressure on Google's profit margins.

"Benchmark scoring involution ( Benchmaxxing )"

Industry experts suggest that Google may be suffering from a common problem in the entire industry: overemphasis on benchmark test scores, which is what is known as the "benchmark scoring involution" phenomenon. Engineers tend to devote more effort to achieving impressive test scores rather than developing products that can actually perform their tasks well. This tendency is common in various artificial intelligence laboratories, as customers often rely on benchmark scores to evaluate the quality of models. Two individuals familiar with this model indicate that Gemini 4 seems to be affected by this issue as well.

Edwin Chen ( Edwin Chen ), the founder of the artificial intelligence startup Surge AI, proposed that if one relies solely on benchmark tests, laboratories will tend to optimize the code capabilities of models specifically for certain programming languages, rather than developing applications with a good user experience and well-designed functionality.

"Let me give you an example: 'Yes, my child SAT got very high scores on the exam.' However, high scores for SAT do not equate to actual abilities in real life," said Edwin Chen. "This is a problem with great potential harm."

Insiders also pointed out that Gemini 4 still possesses advantages. This model is adept at processing input in forms other than text, such as extracting metadata from videos; moreover, it performs exceptionally well in terms of security protection and network security, with the output content being clear and naturally expressive.

The sense of frustration within Google is quite evident. Previously, in a report by Bloomberg that interviewed researchers, they criticized the company for having a large and complex organizational structure, attempting to forcibly integrate artificial intelligence technology into almost all of Google's product lines; the constant changes in task objectives and priority adjustments make it difficult for the company to implement a coherent and unified development strategy.

At the same time, a group of top artificial intelligence researchers have left Google one after another, including legendary engineer Jeff Dean ( Jeff Dean ), Nobel Prize winner John Ng ( John Jumper ), and Noam Shazell ( Noam Shazeer ); the underlying technology he contributed to is a crucial cornerstone of this current wave of artificial intelligence innovation. In August of this year, Demis Hassabis ( Demis Hassabis ), who has long been in charge of Google's artificial intelligence research, took on the role of chairman, handing over the day-to-day operations of the DeepMind department to his long-time deputy, Coray Cardukcuoğlu.

Google announced on Wednesday that Gemini is designed for software engineering, finance, law, and cybersecurity fields, where tasks are often lengthy and complex. Compared to previous versions, this model can generate longer responses, with a single output reaching up to one million words (token), which is approximately equivalent to seven hundred and fifty thousand words.

While Google was striving to catch up, its competitors continued to move forward at high speed. Despite several incidents of AI intelligents entering external organizations, some believed that the development speed of certain cutting-edge models might slow down. Earlier this month, Meta platform company released Muse, a AI intelligent product that is claimed to be capable of independently handling daily tasks such as online shopping and appointments. Shortly after its launch, this app quickly rose to the top of the app download charts.

Tip
$0
Like
0
Save
0
Views 27
CoinMeta reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Mechanic 14 Ultra Thin and Light Notebook X7 358H model available: 1.1kg, retail price of 7224.15 yuan
Mechanics have listed a 14-inch Ultra ultrabook on an e-commerce platform, equipped with an Intel Core Ultra X7 358H processor, 32GB of memory, and 1TB of SSD. The initial price is 8499 yuan, and after national subsidies, the actual price is 7224.15 yuan. It will be available for purchase on October 8th.
The Block
·2026-10-01 11:16:34
7
Measures to deal with AI crawlers: Reddit will suspend RSS subscriptions and terminate public API access
Reddit indicates that support for the RSS subscription source will be discontinued starting from November 13th, and public access to API is scheduled to be closed in March 2027. The company states that these adjustments are necessary to address large-scale data crawling and automated abuse, and they will also affect moderators, third-party applications, research tools, and some AI products.
The Block
·2026-10-01 11:16:33
8
Driven by the growing demand for 4nm process substrates, Samsung's wafer foundry is reported to have seen a year-on-year reduction in losses of 41.8% in 2026
According to reports from various media compiled by JIBANG Consulting, driven by the growing demand for 4nm process HBM4 substrates, Samsung Electronics' wafer foundry and System LSI business are expected to see a 42% year-on-year reduction in consolidated operating losses in 2026. The report also states that the revenue from HBM4 is expected to account for more than 60% of Samsung's total HBM revenue in the second half of the year, and this will drive an increase in orders for advanced nodes and capacity expansion.
The Block
·2026-10-01 11:16:31
8
CNBC Daily Market Opening: Potential Diesel Export Ban, Pentagon's War Preparation Plans, and Google's Latest AI Model
U.S. President Trump said he is still "considering" a ban on diesel exports, raising concerns that energy prices may rise further; meanwhile, the market continues to focus on the Middle East situation, the Pentagon's efforts to build the capabilities needed for future wars, and Google has released its latest AI model Gemini 4 Argon.
CNBC
·2026-10-01 10:46:21
20
Trump claims South Korea's $200 billion investment plan will boost Alaska LNG project
The United States and South Korea announce South Korea's "historic" investment in the U.S. Trump says the new $200 billion investment will change the U.S. "for generations." The investments include nuclear power plants, natural gas power generation facilities in Texas, as well as the potentially revived Alaska LNG project, although whether the project will ultimately proceed still depends on commercial feasibility.
CNBC
·2026-10-01 10:46:19
23
View More