You token maxxin’ too?

Written by

in

It is unsurprising to see organizations scrambling to control AI costs. I am spoiled by subsidized tokens. On Claude I select Opus 4.8, set thinking to MAX, and run my prompt … because I can. On Claude Max or GPT 5.5 Pro, the limits are so generous you feel like by not doing it, you are wasting quota. 10x FTW.

In the current token economy I do not see a rationale to downgrade my model for a prompt to save tokens. Default and go. Times are changing. AI companies are pushing organizations into usage based billing. GitHub copilot changed its policy last month. And people were pissed. Makes sense. You started at $50-100 a month with subsidized usage. Subsidy is gone, and your bill is $2000-$5000.

Will AI companies push everyone to usage based billing? Unlikely but possible that compute is less generous to the point that it’s noticeable. Similar to what we get with Claude Pro vs Max. To save on cost, will people want to invest in local LLMs? Microsoft is reportedly looking into Deepseek to reduce token costs. I think what’s next is a few things.

  1. We get better with token optimization. Select a GPT 4, when 5.5 isn’t required. Start pinching pennies on our prompts and use the dials when appropriate.
  2. Local models for appropriate tasks. While local models that code need beefy systems, professional office work can be serviceable with models < 16B parameters. My experience is that context switching models is not intuitive.
  3. Maybe this is solved by a system like OpenRouter where all models are accessible and can be programmatically chosen depending on the task. Because to expect me to hit a drop-down select every time I prompt or mid-context, is just not happening.
  4. Frontier models keep getting better to the point that older models become more affordable and “fine” for everyday coding. The problem here is that compute is limited. If we look at token costs for Claude, over the past year, the cost per token has not decreased for older Sonnet or Opus versions – but between models there is a notable cost difference. But compute is compute, the cost premium is on the model and effort, so affordability is gained via product choice, not a compute matter.
  5. Local models for everything. Unlikely with the current token economics of memory. Increasing unified memory from 16 GB to 48 GB on a MBP puts a $2K laptop at $4K. This is the bare minimum needed to run a 32B parameter model. And you will likely still swap. 64 GB MBP forces you into Ultra territory at $5K. The trade-off is real for businesses where a $5K laptop may save under usage based billing scenarios.

In the consumer world, frontier AI companies hungry for fresh data are fighting for users. The cost subsidies will continue. There is talk of a price war now that models are mostly commodities. Google kicked it off by making Gemini AI Pro half the cost while giving a lot of Google Premium Perks. It made me take a hard look at my setup. But the Claude harness is too good right now.

If I was forced into usage based billing, I would have to look at my setup and make hard decisions for how to optimize. I would likely start by switching Frontier companies. I often need to go over to GPT or Gemini when I hit my usage on Claude Max. And.. for most of what I do it is fairly interchangeable because they’ve all copied Claude’s harness which improves user choice.