SØNDAG
2026-09-13

Too many projects, too many ideas, too few hours — one learning a day anyway

The Inference Paradox Is Just My Token Bill With a Name

Cheaper models never made my bill smaller. Every time per-token prices dropped, I let the agent loop one more time, add a verification pass, chain a second model. Gartner calls this the inference paradox and predicts a fivefold cost increase per agentic workflow by 2028. That tracks with what I see: a chatbot answers once, an agent reasons, negotiates, and second-guesses itself, and each of those is billable.

So I’m treating routing as infrastructure, not an optimization I’ll get to later. Cheap model for the boring hops, expensive one only where judgment matters, hard ceilings per workflow. Gartner’s phrase for skipping this is “unlimited costs,” and their other number — over 40 percent of planned or deployed agents killed by 2027 — reads like the bill arriving.


The story — Gartner forecasts that inference costs per agent-based workflow will rise fivefold by 2028, despite falling AI model prices. Analyst Will Sommer attributes this to an “inference paradox”: cheaper tokens make complex workflows viable, driving up consumption and total cost. He says no reliable per-unit cost model is in sight, and that companies need optimized inference tiering, routing, and orchestration across multiple models to avoid “unlimited costs.” (Source)