Julia's dad called this week and asked two questions: "What's the deal with companies burning through millions of tokens?" and "How do you decide which model to use?" Chances are if your dad is asking about token budgets and AI models, you might be too.

Let's get into it.

User Error

There's been a wave of headlines about companies pulling back on AI spending after employees burned through massive token budgets. Which is unsurprising, because earlier this year “USE AI!!” was a sudden company mandate, but without telling people why, how, or what to actually use it for. That's like giving every employee a company credit card with no spending guidelines and being shocked when the bill comes in.

The issue likely isn't that people are using AI too much. It's that they're probably not thinking about what model to use and why. The model you use to summarize a meeting should be different from the one you use to write a client report, which should be different from the one you use to classify a thousand data points.

Not All Models Are Created Equal

If you've seen names like Opus, Sonnet, Haiku, Fable, GPT Sol, GPT Terra, and wondered what the difference is, you're not alone. Keeping track feels like trying to remember everyone at freshman orientation.

Just like people, each model has a personality… and some of them should not be trusted with certain tasks.

Here's how we explained it to Julia's dad:

Lightweight models (Haiku, GPT-5.6 Luna, GPT-5.4 mini/nano) are your interns. They’re fast, inexpensive (but fairly compensated, of course), and great at straightforward, high-volume work: classifying data, extracting information, summarizing documents, reformatting text, and handling repetitive tasks. You don’t need your most expensive model doing this work. You need something quick, reliable, and cheap enough to use at scale.

Mid-tier models (Sonnet, GPT-5.6 Terra) are your senior employees. This is where the bulk of serious day-to-day work should live: writing, analysis, research, coding, working across documents, using tools, and solving problems that require judgment. They’re capable enough to handle complicated work without paying for maximum intelligence every time.

Heavyweight models (Opus, Fable, GPT-5.6 Sol) are your senior partners, the C-suite. You bring them in when the problem is difficult, ambiguous, or important: a client-facing report, complex strategy, difficult reasoning, high-stakes analysis, or a final review where quality and nuance really matter. 

And increasingly, the smartest setup is to use these models together. Let the stronger model act like a manager: break a complicated problem into pieces, delegate simpler work to cheaper models, and then review and synthesize the results.

Here’s the basic rule of thumb: Use the cheapest model that can reliably do the job, and move up the ladder as the complexity, ambiguity, or cost of being wrong increases.

Picking The Right Model (In Practice)

We've been fighting with one of our agents this week because she's been, let's say, a bit uncooperative. So we decided to conduct a big audit and planning session to dissect every piece of how she works. Our approach for these types of tasks is typically "Let's throw Fable at it." Fable is the most powerful model, and because this was a complicated review that required real reasoning about architecture, edge cases, and tradeoffs, it made sense. 

But our day-to-day work runs on different models depending on the type of output we’re optimizing for. If we ran Fable on every one of those tasks, we'd be burning money for no reason.

In Claude Code, you can easily switch between what model you’re using. And, if you’re unsure which is the right one, ask your chat.

Before you panic about your AI bill, look at what's actually being used and for what. Chances are, 80% of the tasks people are running could use a cheaper, faster model with zero loss in quality. So staff up your summer interns and save the expensive thinking for the work that actually needs it.

Stay curious,
Julia & Russell

Reply

Avatar

or to participate