
Google is facing reported internal scepticism over how consistently its new Gemini 4 Argon model handles coding outside benchmark tests, even as the company reports leading results on several evaluations.
Bloomberg reported on 30 September that people with direct access to the effort said Gemini 4 performed less well when employees put it to work on some practical tasks, including certain coding jobs. Google disputed that characterisation.
Google unveiled Gemini 4 Argon on 30 September as the top-tier model in its new Gemini 4 generation. The company says Argon is already powering internal workflows, with thousands of employees highlighting strengths in specialised coding, deeper research and writing.
Initial access is being given to a select group of trusted cyber defenders through Google's Fairwind Program. Google says access will expand after further testing, starting with paid API customers and Google AI Ultra subscribers. It has not announced a date for broader public access.
Some Staff Report Uneven Coding Performance
Bloomberg reported that Gemini 4 performed well on commonly used industry benchmarks but did less well when employees used it on some real-world tasks, according to people with direct access to the effort.
People familiar with internal evaluations described its coding abilities as uneven. One person said the model was not particularly adept at front-end design, which shapes the visual and interactive elements of websites and applications.
Bloomberg said there was a range of opinions inside Google. Some employees believed rival models were improving faster, while others believed Gemini 4 had caught up with the leading AI labs.
Google told Bloomberg it would be inaccurate to say Gemini 4 was underperforming in areas such as coding. A Google employee familiar with model development also said there was a 'large consensus' inside the company that Gemini 4 remained at the frontier and denied that the models struggled with messy, real-world coding tasks.
Google Reports Strong Benchmark Results
Google's launch materials report strong results for Gemini 4 Argon. However, the figures are reported by Google and do not by themselves amount to independent validation of the model's broader real-world performance.
Google says Argon scored 77.9% on DeepSWE v1.1, which evaluates long-horizon software engineering tasks, and described the result as a new state of the art. It also reported a 51.3% score on AutomationBench, placing Argon first on the end-to-end business workflow test.
Argon scored 91.7% on LVBench, which measures long-video understanding, and 68% on CWE-bench v1, which evaluates vulnerability remediation. Google said the CWE-bench result tied for first place.
The broader coding comparison is more mixed. Reuters reported that Argon posted better self-reported results than OpenAI's Astra and Anthropic's Opus on several industry benchmarks but remained behind on other measures, including two of the four coding-related benchmarks Google included in its release.
Different benchmarks measure different tasks and use different evaluation methods. A leading result on one test therefore does not establish that a model will outperform competitors across every coding workload.
Google Cites Internal Coding Projects
Google has also published examples of Argon being used on engineering work inside the company.
The company says Argon agents are working on C and C++ migrations to Rust, ranging from tens of thousands of lines in core libraries such as re2 and libgav1 to more than 800,000 lines for the Fuchsia OS Zircon kernel. Google says those large-scale rewrites are undergoing automated and manual auditing, emulation testing and review before production deployment.
In another project, Google says Argon agents started with an existing Rust version of its libgav1 video decoder and replaced 32,000 lines of SIMD code. The company says the resulting decoder ran 2.7 times faster than the Rust version it began with while producing identical video output.
Google also says a team of Argon agents analysed fleet-wide profiling telemetry and applied memory optimisations that freed more than 300 TiB across its data centres once deployed.
These are company-provided examples rather than independent evaluations. They show how Google says Argon is being used internally, but they do not establish how consistently the model will perform across outside developers' projects.
Benchmarks Do Not Settle the Real-World Question
Bloomberg also raised the issue of 'benchmaxxing', an industry term for focusing engineering effort on benchmark scores in ways that may not translate into equivalent gains in practical use. Two people familiar with Argon told Bloomberg the model appeared to be affected by that tendency.
That remains an attributed assessment, not evidence that Google deliberately designed Argon to manipulate or game benchmark results. Google disputes the broader suggestion that the model underperforms in real-world coding.
The launch also follows Google's decision not to release Gemini 3.5 Pro, which CEO Sundar Pichai had previously said was expected in June.
Reuters reported that a Google spokesperson confirmed the company no longer plans to release that model.
For developers, the unresolved question is how consistently Gemini 4 performs on real projects outside Google's own evaluations and internal examples. Wider access will allow customers and independent researchers to test that more directly than any single benchmark can.




