I reach for the expensive model by reflex. Coding, drafting, debugging — Fable 5 or nothing, and I never look at the bill. French-Owen’s numbers made me look: a personalized news site that ran about a dollar per go on Sonnet-class models runs roughly ten cents on gpt-5.6-luna, at ~100 tokens per second.
So I’m splitting my little tools in two. The boat-log summarizer and the feed-triage script don’t need a genius; they need something fast that just handles it. Frontier model stays for work that needs a novel answer. His caveat is the real job, though: harnesses, prompt injection safety, roles and permissions. That’s the part I’d have to build myself.
The story — Calvin French-Owen writes that after weeks with gpt-5.6-luna he finds small models shockingly capable and fast, with research across thousands of emails costing tens of cents. He argues token costs have kept consumer AI companies rare, and that his personalized-news eval dropped from ~$1 on Sonnet-class models to ~$0.10. Citing co-founder Peter, he splits work into “IQ 180” problems and “token spewer” responsiveness — and expects demand for cheap, good-enough models to take off. (Source)