Google's first frontier model in seven months matches Astra on capability but burns twice the tokens per task—turning a lower rate card into a higher total cost.
My Argon invoice for one afternoon of evals came back at almost double what the same work costs me on Astra. The benchmark tables I’d read that morning said the two were a dead heat. The billing meter had a different opinion, and the meter is the one that signs my expense report.
Google shipped Gemini 4 Argon yesterday—its first frontier model in seven months—into a limited preview for what the blog post calls “select testers.” I’m one of them, which mostly means I got an API key a few days early and a Slack channel where a DeepMind PM answers questions between the hours of never and rarely. I spent the first evening the way I spend every model launch evening: I pointed my existing eval harness at it and went to make coffee.
The first hour: it’s genuinely good
The capability story is the easy part, and it’s real. On the independent testing The Decoder published, Argon matches OpenAI and Anthropic without pulling ahead. Read that sentence twice, because the marketing will not. It matches GPT-6 Astra across most of the reasoning and coding suites and trails Claude Opus 5.5—not by a canyon, but consistently enough that if you’re optimizing purely for correctness on hard tasks, Opus is still the one to beat.
So think of Argon as the new restaurant that’s finally as good as the two places you already love. That’s an achievement after seven quiet months. It is not a reason to cancel your other reservations, and anyone telling you it’s the new leader is reading the press release, not the test results.
The headline feature is the output window. Argon will emit up to 1 million tokens in a single response, up from the old 64K ceiling. That’s not a rounding-up. That’s the difference between a model that can draft a chapter and one that can draft the book, generate a full migration across a repo in one pass, or write out an entire synthetic dataset without you babysitting a continuation loop. More on who that actually helps in a minute.
The second hour: the meter
Here’s the rate card, because it’s the thing everyone screenshots. Intro pricing is $2 per million input tokens and $10 per million output tokens. Those are promotional numbers. They step up to $4 and $20—a flat doubling—once the intro window closes. Google hasn’t published the exact cutover date beyond “introductory,” which in vendor dialect means “until we decide you’re hooked.” Budget as if the $4/$20 numbers are the real ones, because by the time you’re in production, they will be.
Two bucks per million input tokens looks like a steal next to the frontier incumbents. It looked like a steal to me too, right up until I actually read my token counts instead of my rate card.
Argon spends tokens like someone else is paying. On the same tasks where Astra would think for a bit and answer, Argon produced roughly twice the total token consumption per task—mostly in reasoning and intermediate output, the stuff you pay full freight for whether or not you ever see it. This matches what the independent testing found, and it matched my meter to an uncomfortable degree.
This is the whole game, so let me make it concrete. Here’s a reasoning-heavy coding task I run as a smoke test—refactor a gnarly module and explain the trade-offs:
On that single task, here’s what the two meters showed me. Rounded, from my own runs—treat them as the ballpark they are, not gospel:
- Astra: ~3,000 in, ~9,000 out (reasoning + answer). Total ~12K tokens.
- Argon: ~3,000 in, ~19,000 out. Total ~22K tokens.
Now run the cost at Argon’s own intro rates versus the step-up. At $2/$10, that Argon task is roughly $0.19. At the $4/$20 step-up, it’s about $0.39. The Astra equivalent, at frontier list pricing, lands in the same neighborhood as the Argon intro number—and below the Argon step-up number, despite Astra’s higher sticker price per token.
That’s the trap in one line: a 50% lower price per token, erased by 2× the tokens per task. You didn’t get a discount. You got a coupon for a bigger meal.
Per-token pricing is the sticker on the gas pump. Tokens-per-task is your car’s actual fuel economy. Nobody brags about cheap gas while driving a truck that does nine miles to the gallon.
The next morning: where the inefficiency stops mattering
I almost wrote Argon off after that first night. That would’ve been the wrong call, and here’s the part I got wrong before I got it right.
The token burn hurts most where you run the same bounded task millions of times—agentic loops, tool-calling steps, RAG answers over retrieved chunks. In those workloads every task is small, you run a torrent of them, and 2× consumption on each one compounds straight into your monthly bill. For agent swarms and high-volume RAG, Argon’s efficiency profile is a genuine liability. Don’t talk yourself out of that math because the per-token number is pretty.
But flip the workload and the picture inverts. The 1M output window is real leverage for jobs where one long generation replaces a fragile chain of stitched-together calls:
- Long-form generation in a single pass. Full reports, book-length drafts, entire code migrations. The overhead you’d normally pay in re-feeding context across a continuation loop can dwarf Argon’s per-task surcharge. One big call beats twelve small ones with twelve redundant context payloads.
- Batch processing with no latency pressure. If you’re running overnight and Google’s batch tier applies its usual discount, the per-task inefficiency gets absorbed into a cheaper lane. You’re trading wall-clock time you don’t need for a rate you do.
- Deep single-shot reasoning where the “wasteful” extra tokens are actually buying you the correctness, and re-prompting Astra three times to get there would cost more anyway.
So the decision rule I landed on: if your unit of work is small and repeated, Argon’s token appetite punishes you and you should stay where you are. If your unit of work is large and singular, the 1M ceiling plus the low intro rate can genuinely come out ahead—provided you’ve modeled the step-up to $4/$20, not the honeymoon rate.
What I’d tell you before you turn it on
Don’t benchmark on correctness alone. Add a column to your eval sheet for total_token_count per task and compute cost-per-correct-answer, not cost-per-token. That single column is the difference between a model that looks cheap and a model that is cheap, and it’s the column every rate-card comparison conveniently omits.
Treat the intro pricing as a timer you can’t see. Build your cost model on $4/$20 so the step-up is a non-event instead of a budget meeting.
And on access: this is select preview, not GA. There’s no general-availability date I’d stake anything on—DeepMind hasn’t published one, and “select testers” today is not a capacity promise for your production traffic next quarter. If a roadmap depends on Argon being generally available on a specific date, that roadmap currently depends on a date that does not exist.
Argon caught up. That’s the honest, slightly boring truth under the launch noise. It caught Astra, it didn’t catch Opus, and it brought a 1M-token output window that’s a real gift for the right workload. Just read the whole meter before you fall for the sticker—the token you don’t see is still a token you pay for.
