In “Opus 5.5 Is The Claude I Missed. One Task Cost Me 50 Cents.”, Nate B. Jones puts Anthropic’s Claude Opus 5.5 through a real job: turning his beanie-and-glasses logo into a buildable 514-piece Lego model with instructions and a parts list. He uses it to argue that the number that matters for a frontier model is not the price per token but the cost per finished task, and that Opus 5.5 brings back the steerable writing many users missed after Opus 4.6.
TLDW
Opus 5.5 turned a logo into a 514-piece Lego model with a 63-page, 58-step instruction booklet, a model file, a BrickLink wanted list, a geometry check file and an LDraw file, using 89 million tokens (at least $44 at API prices) but only about 1% of a roughly $50-a-week subscription, or about 50 cents. Opus 5.5 lists at $4 per million input tokens and $20 per million output, 20% below Opus 5 and 60% below Fable 5.1, and Anthropic says typical workloads cost about 40% less thanks to lower prices and fewer tokens per task, echoed by GitHub, Lovable and Spotify. Jones argues visual work is increasingly done through code like Three.js, which makes surgical revisions possible; that writing is much easier to steer after months of complaints about argumentative Claude models ignoring instructions; that long overnight runs need clear stop conditions and a definition of done because the model keeps pushing; that Anthropic using Claude to write more than 80% of its merged code helps explain fast, feedback-driven releases about 18 days apart; and that everyone should build a personal benchmark by re-running real past tasks on each new model and tracking tokens, usage, failures and corrections.
Thoughts
The most useful idea here is the shift from price per token to cost per finished task. A per-token price says nothing about how many times the model rereads files, how many attempts a change takes, or how often you have to repeat an instruction. Jones’ Lego test is a good illustration because the job is only done when the pieces actually connect, the parts list matches the model and the instructions agree with both. The 89 million tokens and 1% of weekly usage figures are one person’s numbers on one task, and he says so, but the method holds up better than any leaderboard: measure the whole job, not a nice-looking screenshot.
The pricing detail deserves a careful read. Anthropic’s claim that typical workloads cost about 40% less combines a 20% list-price cut with fewer tokens per task, and Jones is right that those are different things. Input, output, cache reads and context each cost differently, so a 40% saving on a bill does not mean a task used 40% fewer tokens. If you want to know what a new model is worth to you, you have to look at all of those buckets, the API bill and, on a subscription, how much of your allowance it used. Most people skip that work and end up arguing over benchmark screenshots.
The “using code to do visual work” point may be the most important one for the future. A Lego model or 3D scene built from code, whether that is Three.js or an LDraw file, is not a finished image. It is a structure the model can go back into and change precisely, like making the beanie taller while leaving the glasses alone, with the parts list and instructions updating to match. That is what makes iteration cheap. It also explains why coding ability keeps showing up as the base skill under everything else the model does, and why the models can now produce tasteful visuals from simple prompts.
The writing section is the emotional heart of the video, and his definition of good AI writing is worth keeping: a good revision preserves your intent. Making a paragraph shorter by removing the complication that made it worth reading, or warmer by softening a decision you wanted stated firmly, sounds polished but makes the writing worse. Jones connects this to public frustration with recent Claude models, including Bram Cohen’s blunt essay on Claude becoming argumentative, and says Opus 5.5 finally feels like Opus 4.6 again, with much more capability behind it. That is subjective, but it matches Anthropic naming clearer writing and better adherence to writing instructions as a headline fix.
The back half gives the practical advice. Opus 5.5 loves to keep going, so long unattended jobs need clear stop conditions and a definition of done, or its persistence turns into wasted tokens. Stating the piece count and complexity up front saved him tokens because the model did not have to decide for itself when it was good enough. Most usefully, his closing challenge is to build a personal benchmark from your own real tasks, run it on every new release with the same prompt and starting point, and record the failures as well as the wins. With releases now arriving roughly every 18 days, a repeatable test like that is the only way to tell a real improvement from marketing, and it is how he decided to move a 90-plus-element design job to Opus 5.5.
Key Takeaways
- Opus 5.5 turned Jones’ beanie-and-glasses logo into a 514-piece Lego model, with the pieces assembling in an animation and the finished model matching the logo.
- The output included a 63-page instruction booklet with 58 steps, a model file, a parts list, a BrickLink wanted list, a check file for geometric brick connections, and an LDR (LDraw) file confirming the pieces connect.
- The job used 89 million tokens, which would have cost at least $44 at API pricing, roughly a quarter of a $200 monthly subscription.
- On his subscription it used about 1% of his weekly allowance. At roughly $50 a week, that task cost him about 50 cents.
- The true cost of a task includes how often the model rereads files, how many attempts it needs, how often you re-explain, and how much it churns, not just the token price.
- Models are spiky, so your efficiency gains depend on your tasks, not someone else’s.
- Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, 20% below Opus 5.
- That is 60% below Fable 5.1, at $10 input and $50 output per million, while matching much of what Fable does on ordinary tasks.
- Anthropic says typical workloads cost about 40% less, from lower prices plus the model using fewer tokens to finish.
- A 40% lower bill does not mean 40% fewer tokens. Input, output, cache and context are priced differently, so check every bucket.
- In Anthropic’s announcement, GitHub and Lovable describe Opus 5.5 using fewer steps, and Spotify describes doing the same tasks cheaper and faster.
- Efficiency matters because it leaves room in your usage window to iterate and improve the result together.
- Do not count tokens too early. A great-looking frame or screenshot is not a finished job, so measure the whole task.
- Claude is increasingly good at using code tools like Three.js to produce tasteful 3D graphics and animations from simple prompts.
- Scenes built from code can be revised surgically, changing one object or movement while keeping the rest.
- Jones says Opus 4.6 was when Claude felt at the top of its game, and that 4.7, 4.8 and Fable made getting the output he wanted harder.
- Public frustration included Bram Cohen’s essay “Why is Claude turning into an asshole?” and GitHub issues about writing instructions being acknowledged and then ignored.
- Anthropic’s 5.5 announcement names communication feedback about Opus 5 as a key issue and lists clearer writing and better instruction adherence as top fixes.
- Good AI writing preserves your intent. Bad edits cut the complication that made an argument worth reading, soften a firm decision, or remove uncertainty you needed to express.
- Writing should be judged by the quality of the result, not by how it was made.
- Opus 5.5 is very steerable: you can say “keep the uncertainty, make this part easier to understand” and get it right the first time.
- In Anthropic’s announcement, Cleo’s Sean Hanks describes Opus 5.5 working unattended for 18 hours on an engineering task across six repos.
- Long overnight runs work best with clear stop conditions and a precise definition of done, because Opus 5.5 keeps pushing itself.
- Specifying the piece count and complexity up front saved tokens by stopping the model from deciding on its own that it was not done yet.
- Anthropic says that as of May, Claude authored more than 80% of the code merged into its codebase. OpenAI describes researchers using agents to write code and run experiments.
- AI helping to build AI shortens the time between user feedback and fixes. Jones says major model releases now come roughly every 18 days.
- A 500-piece Lego build would have broken his Lego benchmark two or three months earlier.
- Users shape models by choosing which ones get their money and their next assignment. The ergonomics of frontier models matter a lot.
- To measure a new model, re-run a task you already did with the same prompt and starting point, then track total tokens, API-equivalent cost and subscription usage.
- Track the failed attempts, corrections and how much you had to step in, and keep the failures as well as the successes.
- After seeing its visual taste, Jones moved a 90-plus-element design job to Opus 5.5 and completed it with one prompt without using much of his weekly budget.
- Build a real personal testing kit from your own work, not someone else’s Lego benchmark, and run it on every release.
Detailed Summary
The 50-cent Lego build
Jones opens with his test: turning his logo, a beanie over a pair of glasses, into a Lego model. Opus 5.5 produced a 514-piece build with a clear column holding the glasses and hat above a base, an animation of the pieces coming together, a 63-page booklet with 58 steps that even flags pieces that attach from underneath, a model file, a parts list, a BrickLink wanted list, a file that checks geometric connections between bricks and an LDraw file. It used 89 million tokens, which he says would cost at least $44 at API prices, but only about 1% of his weekly usage. His plan works out to around $50 a week, so the task cost him about 50 cents.
Cost per task, not cost per token
The token price is only part of the cost. What matters is how many times the model reads the files, how many attempts it takes to make a change, how often you have to explain again and how much it churns. A model that does the job efficiently brings the cost of getting work done way down. He admits this gets squishy because tasks differ and models are spiky, which is why he wants viewers to measure their own gains.
The pricing
Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, 20% below Opus 5 and 60% below Fable 5.1 at $10 and $50. Anthropic says typical workloads cost about 40% less on Opus 5.5 through lower prices and fewer tokens per task. Jones stresses that this does not mean each task used 40% fewer tokens, because input, output, cache and context are all priced differently. He finds the 40% figure plausible given his own test, reports in Anthropic’s announcement from GitHub, Lovable and Spotify, and users on X.
Measure the whole job
Every retry during the Lego build meant more work for Claude, so getting the design right sooner and handling changes cleanly leaves more room to improve the build together. Bigger scenes, like someone’s model of lower Manhattan, will use more tokens. His warning is to stop counting tokens too early: a great screenshot is not a finished task, and real work needs to be measured as a whole job.
Visual work through code
Tools like Three.js let a model build 3D graphics in a browser with objects, materials, lights and cameras. Anthropic’s models have become very good at using these tools from simple prompts to produce tasteful results. Because the scene is code, Claude can surgically change specific objects and movements, such as making the beanie taller while keeping the glasses, or slowing the assembly and adding steps, with the bricks, parts list and instructions staying in agreement.
The Claude he missed
Jones says Opus 4.6 felt like Anthropic at the top of its game for writing and coding, and that through 4.7, 4.8 and Fable getting the work he wanted became harder. He points to community frustration, including Bram Cohen’s essay “Why is Claude turning into an asshole?” and a GitHub issue describing writing instructions being acknowledged then ignored. With Anthropic growing quickly and reportedly pursuing an IPO, a model that makes people spend the afternoon correcting it will lose users. Anthropic’s 5.5 announcement names communication feedback about Opus 5 as a key issue and lists clearer writing and better instruction adherence as top fixes. To Jones, 5.5 feels much closer to 4.6 with far more intelligence behind it.
What good AI writing means
Not all AI writing is slop. A good revision preserves your intent. Asked to simplify, a model can cut the complication that made the argument worth reading. Asked to warm up the tone, it can soften a decision you wanted stated firmly, or make a sentence more confident than you meant. Those edits sound polished but make the writing worse. Jones judges writing by the quality of the result, not how it was made, and finds Opus 5.5 easy to steer: tell it to keep the uncertainty and make a section easier to read, and the fix comes back right the first time.
Long runs need a definition of done
Anthropic’s announcement quotes Cleo’s Sean Hanks leaving Opus 5.5 on an engineering task across six repos for 18 hours. Jones ran two or three overnight tasks himself and found the model does best with clear stop conditions and a precise definition of done. Opus 5.5 loves to keep going, and that persistence should be attached to the job you gave it. For the Lego build, stating the piece count and complexity at the start saved tokens because the model did not have to decide what good enough meant and then keep pushing past it.
AI building AI
Anthropic says Claude authored more than 80% of the code merged into its codebase as of May, and OpenAI describes researchers using agents to write code, run experiments and troubleshoot. Jones sees this as the reason feedback turns into fixes so quickly: users point out what makes a model hard to work with, and the people fixing it have capable AI tools to investigate and test changes. Major releases now land roughly every 18 days. He says a 500-piece build would have broken his Lego benchmark two or three months ago, and that he uses whatever model works rather than picking favorites.
Build your own benchmark
To find out what a release is worth to you, take the files you normally work on (spreadsheets, docs, a code repo), re-run a task you already did with the same prompt and starting point, and track total tokens, the API-equivalent cost and your subscription usage. Record failed attempts, corrections and how much you had to step in, and keep the failures as well as the wins. The result must meet at least the standard of your previous model. After seeing Opus 5.5’s visual taste, Jones moved a 90-plus-element design job to it and got it done with one prompt. His challenge is to build a real testing kit from your own work and re-measure on every release.
Notable Quotes
“Opus 5.5 makes me want to give Claude more work. The writing is so much easier to steer in this model.”
Nate B. Jones, opening the review
“Was 89 million tokens. Do you want to guess how much of my claw usage that took up with Opus 5.5? This is all Opus 5.5. 1% of my weekly usage for that.”
Nate B. Jones, on the 514-piece Lego build
“So, when a model gets the job done efficiently, the price of what you’re getting done can go way, way, way down.”
Nate B. Jones, on cost per task
“We need to stop counting our tokens too early. If you’ve got a greatl looking frame, if you’ve got a greatl looking screenshot, that’s wonderful. You still have work to do.”
Nate B. Jones, on measuring the whole job
“Good AI writing preserves your intent. A good revision has to preserve your intent.”
Nate B. Jones, on the line between AI slop and useful writing
“This things just loves to go and go and go, right? And I love that persistence and I want it attached to the job I gave it.”
Nate B. Jones, on stop conditions for long runs
“These models push themselves. And so it’s up to us to give them a world that they can operate within and be efficient.”
Nate B. Jones, on defining done
“Ergonomics of Frontier models matter a lot. They matter a ton.”
Nate B. Jones, on why user experience drives which models win
Watch the full Opus 5.5 review here.
Related Reading
- Anthropic pricing current per-token rates for Claude models, the starting point for any cost-per-task calculation.
- Three.js the JavaScript 3D library behind much of the code-driven visual work described in the video.
- LDraw the open standard for Lego CAD models used by the LDR file Opus 5.5 generated.
- BrickLink the Lego parts marketplace where a wanted list like the one in the video can be ordered.