insights

When cheap tokens get expensive

Falling token prices can make unit inference cheaper while agentic workflows still become more token-hungry. The useful question is whether workflow cost, not just model price, is improving.

AI economics

Falling token prices are real. But lower unit prices do not automatically mean AI work gets proportionally cheaper once workflows become multi-step, retrieval-heavy, and loop-driven.

Four-panel chart showing token prices, token usage, orchestration inflation, and net task cost after cache

The token-economics example in the Applied Analysis repo was built to answer a simple question: are falling token prices reducing cost per task, or are agents and orchestration increasing token usage enough to offset part of the savings?

The two signals

The analysis compares two signals that become more useful together than separately:

If both fall, the workflow is getting cheaper and lighter. If both rise, richer workflows are also getting more expensive. But if prices fall while usage rises, the signals conflict. That is the case this example is designed to illustrate.

The unit economics improve. The workflow economics improve more slowly.

What the example shows

Estimated blended token list prices decline sharply across the period while estimated workflow token usage per task rises materially. That means the market is getting cheaper at the token level even as agentic patterns push more tokens through each completed task.

The effect is not trivial. Planner, retriever, tool, executor, and critic loops do not just add a small amount of extra prompt context. They can multiply total token demand relative to a simpler single-pass baseline.

Practical implication: evaluate agentic workflows at the task and operating-model level, not only at the vendor price-sheet level.

The useful takeaway is not that cheap tokens make AI expensive. It is more specific: lower list prices can still reduce net cost per task while orchestration inflation absorbs a meaningful share of those savings.

If a pilot depends on agents, retrieval, and repeated review loops, the production case should include workflow-level token economics before scale decisions are made.

Source: adapted from the Applied Analysis example repo post “When Cheap Tokens Get Expensive: The Hidden Cost of Agentic Workflows.”