Technology

AI Token Prices Are Falling. I Asked an Engineer Why the Bills Aren’t

xAI Token

 

Cheaper tokens should mean cheaper AI, so why are many invoices telling a different story? Oxagile’s AI Innovation Lead explains what engineering teams are actually paying for. 

Per-token prices have been moving down. Models can do more, and the hardware underneath them keeps getting better. Yet once teams graduate from occasional AI use to agents doing serious engineering work, the monthly bill has a habit of going the other way. 

Part of the explanation is boring: they use more AI. But production work also burns inference in places that are easy to miss. The agent reads context, tries an approach, runs the tests, goes back when they fail, and may make several passes before an engineer sees something worth reviewing. The first answer can be cheap, but getting to one you would ship in production is another matter. 

The token bill also misses the engineer designing the workflow, the expert reviewing the result, the infrastructure around it, and the work required when an agent confidently gets something wrong. 

I talked to Alexey Karankevich, AI Innovation Lead at Oxagile, and asked: if the unit price is falling, why isn’t the engineering bill following it down? 

The token bill is only part of the bill 

Inference is the easiest cost to see because it arrives with a number attached. Human review is harder to count, even when the project depends on it. 

 

That got me wondering, how much does the token bill really tell us about the cost of agentic engineering? 

 

Alexey commented: “Any accounting that counts only tokens is describing part of the cost. Any plan that assumes the human could have been dropped is describing a project that would have failed.” 

 

Sometimes the expensive workflow is expensive for a reason. And there is an easy way to make that bill smaller: review less. 

 

What does a well-engineered agentic workflow look like at the top of the cost range? 

Alexey shared: “It’s not a cheaper way to write code, but a faster way to do work that would otherwise have been deferred indefinitely.  

The review is part of the workflow and part of the cost. If two engineers are challenging what the agent produces before anything gets accepted, you cannot remove their time from the calculation. 

Strip that out to save money, and you may end up paying for a slower failure.” 

What if token prices stop falling? 

There is another assumption hiding inside today’s AI budgets: that inference will be cheaper tomorrow than it is today. Inference still depends on power, compute capacity, and chips, and model providers have their own economics to account for. 

 

How safe is it to build a budget around that assumption? 

 

Alexey pointed out: “There is no guarantee inference prices will keep falling. More compute capacity and new generations of AI chips push costs down, but provider economics pull the other way. 

 

The leading model market is concentrated among a handful of companies, and pressure to reach profitability could eventually show up in pricing. Teams should be careful about building their AI economics around permanently cheaper tokens. 

 

None of that predicts a particular price movement, and there is no value in pretending otherwise. 

 

The practical conclusion is narrower: falling per-token prices are not a budget plan. A roadmap that assumes this year’s consumption at next year’s lower rates is quietly assuming the one variable that has never held still.” 

 

Reliability is bought with reasoning 

One thing I took from my conversation with Alexey is that a rising bill does not necessarily mean a team is getting worse at using AI. They spend more because they raised the bar for what counts as finished. 

 

A first-pass generation is cheap, yet something you would put in front of a client may require more reasoning steps, evaluation loops, self-checking, retries against tests, and second passes over the same code with a different objective. 

 

That stricter definition of “done” consumes more inference. The team with the smallest bill may simply be stopping earlier. 

The workflow is where the money goes 

Several credible approaches for agentic engineering are already in use, supported by tooling mature enough for production work. Adoption, however, still varies significantly between companies and projects. 

 

What methodologies are starting to stick? 

 

Alexey noted: “SDD and SPDD are both strong candidates. The interesting question is whether the industry will ever converge on just one methodology. So far, I see no evidence of that. Agentic engineering may remain much more dependent on the company, codebase, and type of work than conventional software development.” 

 

That leaves teams making consequential choices project by project. How they split tasks, build context, verify output, and involve engineers all affect consumption. 

 

What guides those decisions today? 

 

The expert admitted: “There is no single, widely adopted methodology yet that we’d call ‘The one’. People will tell you there is, and what they have is a set of habits that worked on their last project. I have those too.  

 

The honest position is that everyone is still working out what this discipline looks like, and anyone offering you a finished framework right now is offering you their habits.” 

Сan you budget for this? 

You can, but request volume is a poor basis for a forecast, but where many finance teams begin. 

 

Alexey explained: “Volume is a symptom. What generates it is the workflow, so that is what a budget conversation should be about: what standard of verification a class of work deserves, how much context each task genuinely needs, where a human has to sit in the loop. 

 

Those are engineering decisions with a price attached. And once they are stated explicitly, the spend becomes forecastable in a way that a request count never is.” 

 

A budget set before the workflow is designed is mostly a guess, as opposed to setting a number in advance and treating overruns as a discipline problem. Teams that get this right tend to run a project deliberately, measure what it actually consumed and why, and use that as the basis for the next one. 

 

How much does model choice really change the bill? 

 

Some advice from Alexey: “Model choice still matters. It is simply a smaller lever than what you put in front of the model, and most teams optimize the two in the wrong order.”  

If the sequencing is the hard part, that is usually a conversation worth having before the first project, which is most of what AI consulting is for. 

Read the bill differently 

The current moment is unusual and probably temporary. Consumption is rising, and teams still vary widely in how they structure agentic work. Those differences are large enough to show up in the bill 

Some of that will settle, as practices will get names, tooling will absorb more of the work, and teams will get better at knowing where expensive reasoning is worth paying for.  

For now, though, the cost of AI engineering is a design question wearing a procurement question’s clothing. Which brings it back to the route rather than the fuel. Per-token pricing is only one input, the workflow decides what you ultimately pay.  

Alexey goes much further into the numbers in one more expert interview. One company spent about $165,000 in inference to rewrite half a million lines of code in eleven days. Was that expensive? The interview also covers why two agents using the same model can produce very different bills, what happens when an agent hits its budget, and why cost per successful outcome tells you more than cost per token. 

 

 

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This