The published snapshot
Selected percentages from Google DeepMind’s model page, rechecked October 1, 2026. These are reported results, not measurements made by Argon Fieldnotes.
| Benchmark | Gemini 4 Argon | GPT-6 Astra |
|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% |
| FrontierSWE v2 | 55.0% | 65.5% |
| Terminal-bench 4.0 | 57.4% | 58.2% |
| CWE-bench v1 | 68.0% | 68.0% |
DeepSWE: Argon +3.8 percentage points. FrontierSWE: Astra +10.5. Terminal-bench: Astra +0.8. CWE-bench: displayed tie. Differences are percentage points, not relative percent improvements.
Why this is not a controlled head-to-head test
The evaluation PDF explains how result sources and settings vary. A published score depends on the harness, available tools and benchmark conditions. Small gaps should not be treated as established differences in reliability without uncertainty estimates or a comparable rerun.
The table does not establish latency, account entitlement or document-extraction quality. Published base token rates are compared separately below; total task costs are not independently measured by us.
A useful coding pilot
If repository maintenance is your goal, begin with a reproducible defect and a passing acceptance check. If a terminal workflow is central to your application, include a task requiring the actual tool environment. Keep instructions, starting revisions and budgets equivalent across candidates.
Use the coding prompt and evaluation protocol to record accepted patches and review effort. Test a wider sample before relying on a narrow success. Confirm Argon access and budget repeated attempts using the pricing guide.
Availability, base prices and governance
OpenAI's announcement describes Astra's staged rollout across ChatGPT and API/cloud channels; it does not prove your account is entitled to use it. Argon's initial access remains the limited channel described in its own announcement. OpenAI rollout statement · Argon access.
OpenAI lists Astra pricing starting at $10 input / $50 output per million tokens. Compare those published rates with Argon's announced stages in the price reference, without assuming equal token consumption or total task cost. OpenAI price source.
For a tool-connected pilot, review OpenAI's Astra system card as well as your actual permissions and logging controls. A safety classification is not a general-purpose task score or proof that another model is safer.