Google has announced pricing for its Argon AI model: $2 per million input tokens and $10 per million output tokens. The rates apply to developers building on the platform. The company disclosed the figures without fanfare, but they carry immediate weight for anyone shipping AI features at scale.
What the two rates actually mean
Input tokens are the text you send to the model — prompts, documents, conversation history. Output tokens are what the model generates back. Under Google's structure, output costs five times more than input. That gap isn't new in the industry, but the exact ratio matters because it mirrors the compute burden: generating text is heavier than reading it.
For a developer, the practical effect is simple. A chatbot that ingests long customer histories and returns short replies will spend most of its budget on input. A tool that turns a brief prompt into a long report will spend more on output. The same model, two very different bills.
Why the pricing may push developers toward efficiency
Because output tokens are the expensive side, there's a direct incentive to keep responses tight. Developers can do that by capping response length, tightening prompts, or caching common inputs so they aren't reprocessed. None of these are new tricks, but the price ratio makes them financially urgent rather than optional.
Input-side optimization matters too. If a company sends a 10,000-token document with every request, that's $0.02 per call just to read it — before the model writes a word. At scale, trimming or summarizing input becomes a cost-control strategy.
The pricing model could also shape product design. Teams might favor workflows that do more with less text, or split tasks between cheaper and more expensive calls. A support bot might retrieve a short, relevant snippet instead of pasting an entire help center into the prompt. A coding assistant might ask for a function rather than a full file.
The market angle
Google isn't pricing Argon in a vacuum. Developers comparing platforms look at both the sticker price and the token mix their application actually uses. A model with cheap output can look expensive if your product generates long responses. A model with cheap input can look expensive if your prompts are enormous.
Argon's $2/$10 split creates a clear signal: Google expects most usage to lean on input, and it's pricing output as the premium action. That could attract applications that read a lot and write a little — search, summarization, classification — while making chatty, long-form generators think harder about their architecture.
It also gives competitors a benchmark. If another provider prices output lower, developers with generation-heavy workloads have a reason to switch. If another provider prices input lower, retrieval-heavy workloads may move. The two numbers are now a reference point in those comparisons, whether Google intended that or not.
What developers can do now
There's no word on volume discounts, committed-use pricing, or free tiers for Argon. The disclosed rates are list prices, and they're per million tokens — which means small experiments cost fractions of a cent, while production traffic can add up fast.
For teams already building on Argon, the immediate step is a cost audit. Count the input and output tokens your application uses in a typical session. Multiply by the rates. If output dominates, look at response-length controls and prompt design. If input dominates, look at retrieval and caching. The math is straightforward once you have the token counts.
Google hasn't said whether these prices will change or how they compare with its other models. For now, the numbers are public, and the design pressure they create is already in place.



