TL;DR
Small, cheap models are cutting AI service costs to just a few cents per request, which will fuel a new wave of consumer AI applications.
The author has been trying gpt-5.6-luna recently. This small model is fast, capable, and more importantly, cheap. Running complex research tasks costs only tens of cents, a completely different magnitude from the previous dollar or so.
Cost is the ceiling for AI apps
In the past, every request in a consumer AI app cost real money, making it impossible to acquire users for free or through viral loops like social networks. Now, cheap small models make the cost negligible, and that playbook works again.
Most of a boss's job is 'token spewing'
Chatting with his co-founder, the author finds that 95% of daily work in companies is not the 'IQ 180' creative thinking, but fast responding and pushing things forward across many fronts - 'token spewing' work. This kind of work happens to be the best fit for cheap and good-enough AI.
Cheap and good enough is what people need
So, frontier models will keep their demand, but demand for 'fast/cheap/good-enough' models might be about to take off. Think of the colleagues, vendors, and customers you interact with daily - nine times out of ten, you just want someone who is super responsive and gets things done.
Curated from high-quality sources, with concise summaries and key takeaways.