Redesign the work before you pick the model

Agents have become good enough to do far more work per request, and that is exactly why AI bills keep climbing. Before reaching for a cheaper model, it's worth asking a different question first: which of this work should still exist at all?

Why better agents can raise your bill

Earlier this year models crossed a line. You can now give one tools and information, and it will reliably investigate a problem, notice what's missing, go and get it, and come back with something you can act on. No more breaking every task into tiny questions and copy-pasting answers between steps.

So people use it more. And each request now does far more inside it: it reads more, calls more tools and follows up on what it finds. More runs multiplied by more tokens per run grows fast. Ten times as many runs, each using a hundred times as many tokens, is already a thousand times the original bill.

It's tempting to blame the people using it. But if every ordinary request takes the most expensive route through the most capable model, that's a design decision someone made about the system, not the user's fault.

Agents carrying the old envelope

Think of the old interoffice envelope: a document goes in, each department does its part, crosses out its name and sends it on. Later the envelope became an email, then a ticket, but the shape stayed the same. Someone still summarises what the previous person said and reformats it for the next team.

Now you can put an agent at every one of those desks. It reads the request, reformats it, prepares the handoff and checks the handoff. It looks very modern, but you're still passing the same envelope around the building. You've made each stop faster without asking why it has to visit all those desks.

Start from the result, not the process

Go back to the handful of value streams that actually matter: winning a customer, delivering what you sold, keeping them and getting paid. For each one, start from a great outcome and draw the shortest path to it, even if that path looks nothing like today's.

Take a customer asking for a quote. Today the request might be summarised for sales ops, checked against the account, translated into a product configuration, entered into a pricing system and passed back through everyone for sign-off. What's really required is much smaller: an accurate quote you're allowed to offer, a record of it, and a question back to the customer for anything you can't resolve. The summary that only existed because one team couldn't read another team's system doesn't need to exist anymore.

Steps you remove cost nothing. Zero tokens is the cheapest model there is. And fewer handoffs also means fewer chances for the request to drift as it moves between people and agents.

This is why business owners have to be part of agent design. An engineer can make a step cheaper, but usually doesn't have the authority to say a department no longer needs a document. Otherwise every team ends up with a faster version of its own process, while the customer is still waiting for the envelope to come back.

Match each piece of work to the right model

Once the unnecessary work is gone, what's left isn't all equally hard:

  • Rules. Pricing, approvals and record updates should be done by ordinary software. The model doesn't need to rediscover arithmetic.
  • Routine interpretation. Understanding which product a customer means, or what they haven't told you yet. A smaller or open-weight model with the right data and tools is usually enough.
  • Real exceptions. Conflicting information and genuine edge cases are where the expensive frontier model earns its price.

A classifier in front decides which kind of request has arrived and sends it to the right place. The goal isn't to find the model that wins at everything; it's to find the model that does this particular job well.

Thick harness, thin harness

The setup around a model (its instructions, tools, data, checks and the path it follows) has to change when the model changes. A smaller model running routine work at scale needs a thick harness: a narrow, clearly defined job, with software checking fields, calling the pricing tool and enforcing approvals. A frontier model working on a hard exception needs a thin one: the problem, the evidence, a few general tools and a clear definition of a good result. Forcing it through every step written for a weaker model gets in the way of the intelligence you're paying for.

Evaluate agents like you would people

You can't tell whether a redesigned process works because an agent sounds pleased with itself, or whether a cheaper model is good enough because its output looks roughly right. We review people's work; an agent with real responsibility needs reviewing too, just far more often than once a year.

For the quote, that means checking the price against the pricing system, that the approval happened and that the record changed, and recognising when asking the customer a question was the right next step. Some checks need judgment, so someone who knows the work has to help define them. Writing good evals is quickly becoming one of the most valuable skills to have.

A bigger bill isn't always a bad bill

If AI lets your team serve more customers, get them the right answer sooner and spend less time repairing handoffs, a growing bill can be perfectly reasonable. If it grew because six agents are writing reports nobody reads, that's a different story. Both can look the same on a token dashboard.

Make it easy for people to use agents, and don't penalise them for doing so. Give ordinary work an affordable path, and let the model spend its intelligence where understanding is actually needed.

A good place to start: take one workflow, follow it all the way to the person waiting for the result, and ask why each step has to happen. You may find a step that needs a more capable model. More often, you'll find a cheaper model with good tools can handle it, or that the step doesn't need to exist at all.