Physix Frontier · News Briefing Card (QbitAI · Oct 1, 2026)

Google Releases Gemini 4 Argon, Tops Benchmarks but Limited to Security Teams

KEY FACTS

  • Google released its new flagship model Gemini 4 Argon, which ranks first on multiple evaluation leaderboards.
  • Argon scored 77.9% on DeepSWE v1.1, surpassing Opus 5.5 and GPT-6 Astra.
  • The model is currently available only to vetted cybersecurity teams, with general users unable to access it for now.
  • Argon supports up to one million token output and targets complex workflows in programming, finance, and law.
  • Some Google employees are skeptical of the model's actual coding performance, believing front-end design remains a weak spot.

KEY DATA

77.9%DeepSWE v1.1 Score
68%CWE-bench v1 Score
68.90%Vals Index Score
2 USDPrice per Million Input Tokens

PHYSIX OBSERVATION

Google is using low pricing and long context to claw its way back to the table, but giving the first batch only to security teams shows it still has concerns about abuse risks. Leading on benchmarks does not equal leading in experience, and internal doubts about the front-end weak spot are worth noting. For developers, the cost-performance ratio is a real positive, but whether it can deliver still depends on real-world testing after full public release.

Source: QbitAI report