Opus 5.5 and the real cost of a task
Opus 5.5 is the first Claude release in a while that feels easy to steer again, especially for writing. But the more useful lesson from it isn't about any one model: the price per token tells you very little about what a piece of work actually costs.
Price per token isn't cost per task
What you pay for a job depends on how the model gets there: how many times it reads the files, how many attempts it needs, how often you have to explain the same thing again, and how much it churns before it settles. A model that finishes in fewer steps can make the same work much cheaper, even at a similar rate.
- Published pricing. $4 per million input tokens and $20 per million output tokens: 20% below Opus 5, and 60% below Fable 5.1 at $10 and $50.
- Anthropic's estimate. Typical workloads cost about 40% less than on Opus 5, from the lower prices plus fewer tokens per job. That's not the same as a task using 40% fewer tokens; input, output, cache and context are all priced differently.
- Early reports. Teams quoted at launch describe the model taking fewer steps, and doing the same tasks faster and more cheaply.
Count the whole job
It's tempting to judge a model on the first good-looking result: a render, a screenshot, a passing test. That's usually halfway. Say you ask for a 3D model of a product along with assembly steps and a parts list. The job isn't finished until the steps, the parts and the model all agree, and every round of fixing that costs tokens and your time.
- Measure the full bill. Total tokens, the equivalent API cost, and how much of a subscription it used, from the first prompt to the finished work.
- Count the hand-holding. How often you had to step in, re-explain or correct is part of the cost, even if it doesn't show up on an invoice.
- Prefer work built as code. Visuals made with code (a Three.js scene, say) can be revised precisely: change one object or one step and everything that depends on it stays consistent. Opus 5.5 is notably good at this kind of visual work.
Easier to steer
The release notes name communication as a focus, with clearer writing and better adherence to writing instructions, after months of feedback that earlier versions acknowledged instructions and then ignored them.
Good AI writing preserves your intent. The failures are subtle: making a paragraph simpler by deleting the complication that made the argument worth reading, making the tone warmer by softening a decision you meant to state firmly, or sounding more confident by erasing uncertainty you needed to communicate. Each edit reads as polished and leaves the writing less useful. What matters is whether "keep the uncertainty, make this part easier to follow" comes back right the first time.
Long runs need a definition of done
Opus 5.5 can work unattended for a long time; one launch customer left it on an engineering task across six repositories for 18 hours. That persistence needs edges:
- Set clear stop conditions and a concrete check for when the work is finished.
- Fix the scope up front. If you state the size and level of detail you want, the model doesn't have to decide them itself, decide it hasn't done enough, and keep going. That saves a lot of tokens.
Test it on your own work
Your tasks aren't anyone else's, and models are uneven: one release can be much better at visual work and barely different at research. The only reliable answer comes from your own assignments.
- Pick a task you've already done with a previous model, using the same files, the same prompt and the same setup.
- Run it on the new model and track the whole job: total tokens, the equivalent API cost, and how much of your subscription it used.
- Count the attempts, corrections and how often you had to step in, not just the final result.
- Hold it to at least the standard of the previous model, and keep the failures as well as the successes.
Repeat it with each release and you'll know when something has really improved for the work you do, and where to send it. The same applies in reverse: when a release saves effort, you can spend less or take the extra room and build something more ambitious.
Why releases keep speeding up
Anthropic says Claude wrote more than 80% of the code merged into its codebase as of May, and other labs describe researchers using agents to write code, run experiments and debug. Major model releases now arrive roughly every 18 days. The encouraging side of that loop: feedback from people using these tools can be investigated and fixed faster, and Opus 5.5's writing is a visible example.