Data as of 2026-10-02 · day 3 after launch

Gemini 4 Argon is one-fifth the price of GPT-6 — so why does a single task save only 40%?

Three answers on one screen: how Argon trades blows with GPT-6, what to put in your API call (including the 1M output-token continuation), and who can actually call it today. No signup, no scrolling.

Who can run Argon today — and who can't

Google has not published a model id for the public API, and has not published an opening date. Right now only vetted partners in the Fairwind Program can call it, and resale is not permitted.

1. Fairwind partners (running today) 2. Paid API users / Google AI Ultra 3. Developers · enterprises · consumers (no date)

Gemini 4 Argon vs GPT-6: feature by feature

The same public benchmarks and independent tests. Only cells with a source.

AreaArgonGPT-6 AstraGPT-6.1 Sol
Writing code: who does it better
Long-context refactoringFrontierSWE v2 55.0%65.5%—
Long-horizon engineeringDeepSWE v1.1 77.9%74.1%—
Finding and fixing bugsCWE-bench v1 68.0%68.0%—
Business workflow automationZapier AutomationBench 51.3%41.4%—
Long terminal tasksTerminal-Bench 4.0 57.4%—Opus 5.5 leads at 66.4%
Money: what a call really costs
List price (input / output, per 1M tokens)Argon promo → standard $2/$10
→ $4/$20
$10/$50 $2/$10
Cached inputmatters when you re-feed the same context $0.10$1.00$0.10
Output tokens burned per taskArtificial Analysis test, fewer is cheaper 62,00027,000—
Real cost per tasksame task on the Intelligence Index, promo pricing $1.99$3.26~$0.74
Reliability and overall
Composite Intelligence IndexArtificial Analysis, high setting 535352
Hallucination rate when wrongAA-Omniscience, lower is better 15%51%—

The short version: Argon's input price is 1/5 of Astra's, but it burns 2.3× more output tokens on the same task (62,000 vs 27,000), so real cost per task only falls to about 60% ($1.99 vs $3.26). Astra still wins on long-context refactoring; Argon leads on long-horizon engineering and automation; for the absolute cheapest run, use GPT-6.1 Sol.

The call skeleton: paste it in, change one line, it runs

The model id is not published yet. The fields below follow the OpenAI-compatible format — on launch day you only change MODEL.

curl
# As of 2026-10-02 there is no gemini-4-argon model id in the public API
POST https://generativelanguage.googleapis.com/v1beta/openai/chat/completions
  -H "Authorization: Bearer $GEMINI_API_KEY"
  -H "Content-Type: application/json"
  -d '{
    "model": "gemini-4-argon",
    "messages": [{"role":"user","content":"Refactor this module"}],
    "max_tokens": 1000000,
    "temperature": 0.2,
    "stream": true
  }'
python (handling the 1M output cap with continuation)
import os, requests

MODEL = "gemini-4-argon"          # not published yet, placeholder
BASE  = "https://generativelanguage.googleapis.com/v1beta/openai"

def run(prompt, cont=None):
    body = {"model": MODEL, "messages": [{"role":"user","content":prompt}],
            "max_tokens": 1000000}
    if cont: body["continuation"] = cont   # Long Decode Continuation: resume where it stopped
    r = requests.post(f"{BASE}/chat/completions", json=body,
        headers={"Authorization": f"Bearer {os.environ['GEMINI_API_KEY']}"})
    j = r.json()
    return j["choices"][0]["message"]["content"], j.get("continuation")

Read these three before you set parameters

  • A 1M output cap is not a 1M input context. Google only published the output cap (raised from 64K to 1M). The input window spec is still unpublished — do not design your chunking around 1M.
  • Very long generations get paused; resume with Long Decode Continuation. A single response is cut off mid-way and you have to continue it with follow-up calls. Skip that and you get half a result.
  • Do not copy the Gemini 3.x price sheet. Batch / flex / priority / cache storage are all still unpriced for Argon, and the promo end date is unpublished — budget at the standard $4/$20.