Google's upcoming Gemini 4 model, codenamed Argon, outperformed rival AI models on 12 of 18 benchmarks listed in the company's internal comparison table. The model also features a million-token reply capacity, leads in resistance to hijacking attempts, and will be released to cyber defenders before anyone else — with safety guardrails disabled.
What the comparison table shows
Google's own benchmark table lists 18 tests across categories that typically include reasoning, coding, math, and language understanding. Gemini 4 Argon comes out on top in 12 of them. That's a majority, though not a clean sweep. The remaining six benchmarks presumably went to competing models, though the facts provided don't specify which ones or by what margins.
The benchmark claims come from Google itself, which is worth noting. Company-run comparisons often use favorable testing conditions or select benchmarks that highlight strengths. Independent verification typically follows after release, and that's when the real picture emerges. For now, the only data available is Google's own.
The million-token reply
Gemini 4 can write a million tokens in a single reply. That's the output side, not the input context window — a distinction that matters. Most large language models today can read long documents, but generating a million tokens in one response is a different scale entirely. For context, a million tokens is roughly the length of several full-length novels. Whether that capacity translates into useful, coherent output at that scale is another question, but the capability itself sets Gemini 4 apart from current models on the market.
Practically, that means tasks like generating entire codebases, long-form reports, or extensive documentation in one go become technically possible. Whether they become practical depends on latency, cost, and how well the model maintains coherence over such long outputs.
Hijacking resistance
Among the AI models evaluated, Gemini 4 Argon resists hijacking best. Hijacking in this context refers to prompt injection or jailbreak attempts — users trying to trick the model into ignoring its instructions or producing harmful content. It's a persistent problem for deployed AI systems, and one that security researchers test aggressively.
Google hasn't detailed the specific hijacking tests used or the scoring methodology. But the claim positions Gemini 4 as more robust against adversarial prompts than its competitors, at least according to Google's own evaluation.
Cyber defenders get first access — guardrails off
The most unusual detail: cyber defenders will get access to Gemini 4 first, with guardrails off. That means security professionals — likely those working on defense against cyberattacks — will be able to use the model without the standard safety restrictions that normally prevent certain types of output.
Disabling guardrails for a specific group is a notable decision. It suggests Google wants defenders to have maximum capability when analyzing threats, writing detection rules, or reverse-engineering malware. But it also raises questions about how Google vets who qualifies as a cyber defender and what prevents misuse. The facts don't specify the vetting process or the timeline for broader releases.
Guardrails-off access for a select group is a departure from how most AI companies handle safety. Typically, safety filters apply to all users, and exceptions are rare. Google appears to be treating cyber defense as a special case where the benefits of unrestricted model access outweigh the risks.
What's still unknown
There's no release date for Gemini 4 Argon, no pricing information, and no details on when non-cyber-defenders will get access. The benchmark table is Google's own, and independent results won't arrive until the model is in wider hands. The million-token reply feature and hijacking resistance will both need real-world testing to confirm they hold up outside controlled evaluations.
For now, the key numbers are 12 out of 18, one million tokens, and guardrails off for a select group. The next concrete step is finding out who qualifies as a cyber defender and when everyone else gets a turn.




