US-built AI models often cost less per task than their Chinese counterparts, according to a new analysis. The finding could nudge enterprise buyers away from token-based pricing and toward measuring what a model actually gets done.
That shift matters because most companies still budget for AI by the token—the chunks of text that models process. But tokens don't tell you whether a model finished the job, or how many tries it needed to get there. If a US model charges more per token but completes a task in fewer steps, the total bill can come out lower.
The gap between token price and task cost
The analysis compared models on a per-task basis rather than sticker price per token. It found that US models frequently come out ahead on that metric. The reason isn't necessarily cheaper tokens. It's that task-level accounting captures retries, tool calls, and the cost of a failed attempt—things a per-token price hides.
For an enterprise running thousands of automated workflows a day, those hidden costs add up. A model that needs three attempts to resolve a customer query can be more expensive than one that costs twice as much per token but resolves it on the first try.
Why CFOs are starting to ask different questions
Procurement teams have long used token pricing as a shorthand for AI spend. It's simple and it's visible on the invoice. But it's also a poor proxy for value. The analysis suggests that as AI moves from pilot projects into production systems, finance departments will want metrics that map to business outcomes—resolved tickets, completed transactions, generated code that passes tests.
That's a harder number to get. Vendors don't always publish task-completion rates, and benchmarks rarely reflect a specific company's workload. Still, the direction of travel is clear: the buyer's question is shifting from "what's your price per million tokens?" to "what does it cost to get this done?"
What this means for Chinese model providers
Chinese AI models have gained attention for aggressive token pricing, sometimes undercutting US options by wide margins. But if the per-task math favors US models, that advantage narrows. Chinese providers may need to compete on reliability and task efficiency rather than raw token cost alone.
The analysis doesn't name specific models or vendors, and it doesn't put a dollar figure on the gap. That leaves room for debate. Task definitions vary, and a model that's efficient for one workload can be wasteful for another. A coding assistant and a customer-service bot don't measure success the same way.
Where the numbers still don't add up
Even if the per-task finding holds, switching costs are real. Companies have built pipelines, fine-tuned models, and trained staff around existing providers. A cheaper task cost doesn't automatically justify a migration, especially when the difference is small.
There's also the question of who bears the cost of a failed task. In some deployments, the customer eats the retry. In others, the vendor does. That changes the calculus entirely. The analysis frames the issue as a spending-strategy question, but the answer will depend on contracts and SLAs as much as on model performance.
For now, the takeaway is narrower: token pricing is an incomplete measure, and US models appear competitive when you count the whole job. Enterprises evaluating AI vendors in the next budget cycle will have to decide whether to trust that framing—or keep shopping on the sticker price.




