Skip to content
model-selectioncostworkflowspractitioner

Think Expensive, Execute Cheap

The money-and-quota side of picking an AI model: why defaulting to the biggest one quietly costs you, and the habit that fixes it.

Fabian Mösli Fabian Mösli
· 7 min read · 2026-06-24

Key Takeaways

  • Model tiers are priced in multiples, not increments. On Anthropic's line the jump from the smallest model to the frontier is a clean 10×, and the biggest single step is right at the bottom, where Haiku to Sonnet triples the cost. 'Always use the best' is a budget decision most people make by accident.
  • Think expensive, execute cheap. Use your strongest model for the judgment: strategy, planning, the hard calls. Then hand a detailed brief to a cheaper, faster model to execute. The judgment lives in the plan; execution is mostly instruction-following, which cheap models do well.
  • On a subscription you pay in quota, not dollars. Plans meter usage roughly in proportion to API price, so heavier models drain your weekly allowance faster. The same habit keeps you from hitting a wall and getting locked out until your limit resets.
  • Build a feel by trying models across vendors and settings. Cost-efficient models, several of them non-American and open-weight, are now genuinely capable. Running the same task through a few of them at different effort levels teaches you which model fits which job.
In this guide

Most people pick a model once and then use it for everything. Usually the most capable one they can get their hands on, because why not use the best? At one task a day, that’s fine. Nobody’s budget notices. But at a hundred tasks a day, or across a team, or inside an automation that runs all night, that one reflex gets expensive fast. And most of what it’s paying for is wasted on work that never needed it.

There’s a whole conversation about which model to use that’s really about capability: reasoning versus fast, chat versus agentic, this benchmark versus that one. This isn’t that. This is the other axis, the one people skip: cost. Not because it’s complicated, but because the tools make it easy to ignore until the bill, or the “you’ve hit your limit” message, shows up.

The tiers are multiples apart, not percentages

Most people never check this. The model tiers aren’t priced a few percent apart. They’re priced in multiples of each other.

Take Anthropic’s lineup, because the numbers come out unusually clean. Put Haiku, the small fast one, at a baseline of 1, and the rest of the ladder looks like this:

  • Haiku: 1×
  • Sonnet: 3×
  • Opus: 5×
  • Fable: 10×

(Those are per-token API prices as of June 2026. They drift, so treat the shape as the point, not the decimals.)

Think about that scale for a second. Top to bottom, it’s a full 10× spread. The frontier model costs ten times what the small one does to do the same unit of work. And the single biggest jump is the first one: going from Haiku to Sonnet triples your cost before you’ve even reached the middle of the range.

There’s a catch, though. That ladder blends two things. There’s the per-token price, and there’s how many tokens a model actually burns — bigger models tend to think more, out loud, before they answer. So the real-world gap can run even wider than the sticker price. For a rule of thumb it doesn’t matter. The takeaway holds either way: the gap is large, and “always use the best” is a budget decision most people are making without realising they’re making it.

The habit: think expensive, execute cheap

Once you see the ladder, the habit is obvious:

Use your strongest model to think. Strategy, concepts, planning, the hard judgment calls. Then hand the finished plan to a cheaper, faster model to execute.

Why does that work? Because the expensive part of most knowledge work is the judgment, and the judgment all lives in the plan. Execution — turning a clear, detailed plan into actual output — is mostly instruction-following. And cheaper models are genuinely good at instruction-following. Where they fall down is when they have to make judgment calls of their own. They shine when a stronger model has already made those calls and written them down.

It’s the same move a good manager makes. The senior person sets the brief. The brief is the thing that makes it safe to hand the rest to someone who costs less and works faster. You’re not delegating the thinking. You’re delegating everything after the thinking.

The trick is the handoff

The bridge between the two halves is a document. Get your strong model to write an execution-ready brief, precise enough that a cheaper model can run it without needing to think for itself:

“Turn this into a precise, step-by-step brief that a faster, cheaper model could execute without making judgment calls of its own: [your task].”

Then open a fresh chat on the cheaper model and paste the brief in. That’s the whole technique. Most of the value is in forcing the strong model to make its reasoning explicit, which, conveniently, also makes the work easier for you to check.

How I actually split it

To make this less abstract, here’s roughly how I divide my own work.

The frontier models get the planning and the deep thinking, the parts where I want a real thinking partner that pushes back. Anything where the quality of the judgment decides the quality of everything downstream.

Mid-tier models get the implementation. Once the plan is solid, turning it into a first draft, a structure, a working version — the medium models handle that well, and they’re quick about it.

And the simple stuff, cleaning up copy, reformatting, a quick rewrite, the small office tasks that fill a day, goes to the flash models or Haiku. There is no reason on earth to spend frontier-model money, or frontier-model quota, tidying up an email.

That split isn’t a rule I follow rigidly. It’s a feel I built up by using the models enough to know which one a given task deserves. Which brings me to the actual recommendation, but first, the part people on a subscription tend to miss.

On a subscription you pay in quota, not dollars

If you’re on the API, the ladder shows up directly on your bill. But most people aren’t on the API. They’re on a subscription — a Pro or Max plan, a flat monthly fee — and it’s easy to assume that makes the model choice free. It doesn’t.

Subscriptions don’t give you unlimited use. They give you a budget: a rolling window of a few hours, with a weekly cap on top. And the model you pick sets the burn rate. The heavier the model, the faster it drains that budget. Anthropic meters plan usage roughly in proportion to the API price, so the same 1-3-5-10 ladder applies. Fable drains your allowance about twice as fast as Opus. Opus chews through it faster than Sonnet again.

So on a subscription you don’t get a surprise bill. You get a wall. You run out, and you’re locked out until the window resets — sometimes a few hours, sometimes the rest of the week. “Think expensive, execute cheap” stops being about saving money and becomes about not burning your whole week on Opus by Wednesday afternoon. Plan the hard problem on the big model, run the execution on a cheaper one, and the same allowance stretches several times further.

The cheap end is better than you’d guess, and it isn’t only American

When people picture “the cheaper option,” they picture a worse model. That’s increasingly wrong, and it’s worth knowing, because it changes the maths.

The big three American labs (Anthropic, OpenAI, Google) aren’t the only game. There’s a group of cost-efficient models, several from Chinese labs and several open-weight, that have quietly gotten very good. GLM-5.2 is the one that surprised me: on some design and coding benchmarks it goes toe to toe with the frontier American models, at a fraction of the price. I’m not telling you to move your whole stack onto it. Benchmarks aren’t real life, and there are good reasons, from data to governance, to think carefully about which models you put to work where. But the idea that capable has to mean expensive, or that the only serious models come from three companies in California, is just out of date.

The real advice: go build the feel

If you take one thing from this, make it this: try them. All of them.

Run the same real task through a few different models. Push the effort or thinking setting up and down and watch what changes. Take a problem to Claude, then to Gemini, then to one of the cheaper challengers, and notice where each one is strong and where it falls apart. It costs you an afternoon. What you get back is a feel for which model fits which job, and that feel is worth more than any benchmark table, because it’s calibrated to your work, not someone else’s. You build that judgment by doing the work, not by reading model cards.

The payoff is three things at once. You get better results, because you stop forcing every task through a model that’s wrong for it. You spend less, because you stop overpaying for work that never needed the premium. And you stop getting blocked, because you’re not torching your usage limit on things a smaller model would have handled fine.

The honest part

I’ll be straight with you, because the whole point of this site is signal, not a sales pitch.

There is no reason to use the biggest models for simple office work. None. I stand by every word above.

And frontier models are addictive. If you do real knowledge work, the best model available becomes the one you reach for without thinking, and every other model starts to feel a little dumb by comparison. I felt exactly this during the few days Anthropic pulled Fable 5 from general access, under a US government directive, by Anthropic’s own account. I was suddenly back on Opus, a model that had been the best in the world until just days earlier, and it felt like a downgrade I could feel. Same model I’d been delighted with the week before. Now it seemed slow on the uptake.

So take the advice with that caveat. The discipline is real and it’s right. It’s also something I have to keep choosing, against a pull toward the best model that doesn’t go away. Just knowing the pull is there helps me resist it.

Where this breaks (because it does)

A few honest limits, so you don’t take the rule further than it goes:

  • Don’t cheap out on the thinking. A sloppy plan executed cheaply is just cheap slop, delivered faster. The expensive model has to earn its keep on the part that actually needs judgment.
  • Some execution genuinely needs the strong model. Anything where the “execution” is itself full of judgment calls (nuanced writing, ambiguous edge cases) doesn’t split cleanly. The handoff works best when the execution is close to mechanical.
  • The overhead has to be worth it. For a one-off task, skip all of this and use one good model. The ladder pays off at volume, in automations, and across a team, not on a single email.

The point is to spend where spending buys you something. Whether you’re paying in dollars or in quota, the choice of model is a budget decision. So make it on purpose, and keep the premium for the problems that earn it.

Published: 2026-06-24

Last updated: 2026-07-01

Stay in the loop

Don't miss what's next

I'm curating the best AI tools for professionals. Join the list and I'll reach out when I have something worth sharing.