,

AI Cost per Task: What Token Prices Miss

Jason Fishbein looking at a laptop marked 3-cent model call above an iceberg representing trusted data, system integrations, permissions, business rules, human escalation, and monitoring.

Token prices are down, but you are still spending more and more on AI.

Maybe that is just the price to pay for intelligence.

It sounds backward until you watch what happens inside a real company.

A model gets cheaper. Nobody uses it less. They give it a harder job.

First, it summarizes a document. Then it summarizes every customer call. Then somebody asks whether it can identify risk, update Salesforce, make a recommendation, or make the decision itself. Then somebody from Legal asks why it made that decision.

And suddenly, the three-cent model call has friends.

Trusted data. System integrations. Permissions. Business rules. Human escalation. Monitoring. A small committee trying to decide whether “customer” means the same thing in all six systems.

The model got cheaper. The work got more serious.

Price per token is not the cost of intelligence

Price per token is a useful number. It tells you what the model costs before it touches your business.

It does not tell you what it costs to get a correct outcome after the model has to understand your customers, follow your policies, use your systems, and avoid creating a problem that needs a meeting with “incident” in the title.

That is the difference between the cost of generating an answer and the cost of completing a task.

For an enterprise AI workflow, the fully loaded task cost includes the model and reasoning budget, retrieval, tool calls, orchestration, retries, human review, exception handling, and the cleanup when the system gets something wrong.

That is why a cheaper model is not automatically a cheaper workflow.

A three-cent answer can create a twelve-dollar problem

Take customer service. A low-cost model can draft a polite answer to a customer in seconds.

A useful AI system needs more. It needs the correct order history, the product context, the current policy, any existing exception, the authority to take an action, and a clear rule for when it needs to stop being helpful and bring in a human.

Otherwise, you have not automated customer service.

You have automated the first paragraph.

That distinction matters because the cheap model call can still create an expensive downstream process. If it takes three attempts, triggers a human review, sends the case to the wrong team, or creates an exception somebody has to untangle later, the token bill was never the main bill.

The cost equation needs to follow the work, not the vendor invoice.

Total workflow cost ÷ successfully completed tasks = cost per successful task.

Not cost per chat. Not cost per million tokens. Cost per task completed correctly.

The metric changes the architecture conversation

Once teams measure the cost per successful task, different questions become more important.

Is the cheaper model causing more retries?

Should this request be routed to a stronger model the first time?

Is human review catching real risk, or rubber-stamping routine work?

Is the exception rate telling us the data, policy, or workflow is broken?

Are we building reusable context and integrations, or paying to rebuild the same plumbing for every new use case?

That is where enterprise AI ROI becomes real. The goal is not to make every response as cheap as possible. It is to get the right outcome for the lowest total cost.

Why spending more can still be progress

There is good news hidden inside the rising spend.

When a company moves from using AI to produce text to using AI to produce outcomes, spend should increase before it starts to compound. More meaningful work requires better context, better controls, and better workflow design.

But those investments are reusable. Trusted data products, integrations, evaluations, guardrails, and escalation patterns make the next workflow faster and less expensive to launch.

That is the difference between companies running pilot number 47 and companies building an operating capability.

Intelligence is cheap when it produces text.

It gets expensive when it starts producing outcomes.

And that may be exactly the point.