Video: Building with the Claude 5.5 family: choosing the right model and getting more from every token | Duration: 2920s | Summary: Building with the Claude 5.5 family: choosing the right model and getting more from every token | Chapters: Welcome and Introduction (8.21s), Welcome and Introductions (60.32s), Model Evolution Showcase (105.395s), Cloud 5.5 Features (231.68s), Migration Changes (388.94s), Common Gotchas (523.7s), Prompting Guide Tips (726.495s), Model Selection Strategy (858.235s), Cost Per Task (1076.64s), Cost Optimization Strategies (1354.55s), Handoff to Ben (1596.145s), Model Implementation Playbook (1636.78s), Cost Optimization Strategies (1905.545s), Model Routing Strategies (2146.435s), Moonbase Pizza Case Study (2295.75s), Model Tuning Process (2527.165s), Cost Optimization (2660.61s), Conclusion and Next Steps (2767.175s)
Transcript for "Building with the Claude 5.5 family: choosing the right model and getting more from every token": Hello, everyone, and welcome. Thanks for for joining us today. We're really excited about the Claude five five family and and hope you are too. Opus five five and SONNET five five give you more ways to match the model to the work. Today is about how to make that choice, which model should you run each of your tasks, and how to get more from every token. A bit of of housekeeping before we start, we're recording today's session. We'll be sending the recording and slides to everyone who registered. They'll be available for download. We won't have a live q and a at the end, but our team is answering questions in the chat throughout. So send send any questions you have over to chat, and we'll get to them. To kick things off, Lucas Gonzales from our research team will share how we built the five five models and what we found in testing them. Then Ben Lareberger from Applied AI will show you how to put that into practice. With that, over to you, Lucas. Awesome. Thanks, Ryan, and thanks everybody for joining. Can folks hear me? Can I get some emoji reacts, please, if so? Awesome. Great. Great. Always love the emojis. Well, yeah, welcome, everybody. I'm gonna be chatting about the Cloud five five family, which we are very, very excited about. Like Brian mentioned, my name is Lucas. I'm a product manager, over here in our research org at Anthropic. I mostly work on model launches, the evolves behind them, model positioning strategy, and then, a focus area that I also spend time on is computer use and browser use. I was previously at Google DeepMind, and I'm very, very excited to share with you guys, some more info on the Cloud five five family. So before we get started, I wanted to take us down a trip down memory lane. And I know we all really love, Opus 4.6. That was a really great model. Of course, I love that model as well. But over time, the models have gotten really noticeably better at doing stuff that matters. And so I wanted to show you guys just kind of three slides that were generated all by Claude. The first one by Opus four point six, the next one by Opus five, and the last one by Opus five point five so that we could really feel the difference through, through the generations. And then the rest of the side deck is, of course, made by Claude as well. So the sample that I gave Claude to start building with here is basically just a very simple query, how do large language models work. So this is a slide that Opus 4.6 put together, and you could see, you know, it has some grasp of elements. Like, it's able to use these, like, big text boxes, put them in order together, kinda walk you through, like, four basic steps, and then give you that, like, insight at the bottom. Next with Opus five, you can see we're starting to get a bit more fancy. It includes that bar chart at the bottom. It includes some, elements with different fonts within the boxes, And the text itself is also a bit more understandable. You see where, you could start to understand the concepts that the model is trying to get across. Finally, with Opus 5.5, you could see it just got way more advanced in terms of the icons it uses, the diagrams, how it's getting the idea across. I particularly thought that that third box of mix in content where it's explaining the attention mechanism was just a really nice touch, by the model. And so point is you can see how, you know, throughout the three generations, we've made really big strides in use cases such as, a slide generation for knowledge work. And sometimes it's hard to tell since the models are already so good and since Opus 4.6 was such a great model, but these improvements have really made a big difference. So with that, I will start us off with, what's new in our cloud 5.5 family. What are things you should know, what are things that you may have missed. And then from there, I'll speak a little bit about how I approach thinking about which model to use when, which my co presenter Ben will get into more detail during his section. And then at the end, I'll speak a little bit about cost per task, how I like to think about cost per task and choosing the right model for a task, as well as how you could reduce, the cost that you're paying for getting great work done with Claude. So starting us off with the five five family. So so far, we've released Opus 5.5 and Sonnet 5.5. We're hard at work on a Haiku for you all as well. The main story behind these two models we've released so far has been, really big quality bumps while also simultaneously reducing the cost to use these models and making them faster to use and better communicators. We are particularly happy about four things that, you know, we feel has really improved with these models. First of all, of course, is code capability, which, continues to improve. But especially on terminal bench, we see both, Opus 5.5 and SONNET 5.5 do much, much better than their, predecessors. And this also shows up in some of the capabilities that we have seen going around on socials as well, such as that visual spatial reasoning. And so we see that the model's doing really well at three d asset creation. It's, quite outstanding at generating games. And, of course, this, like, visual reasoning also helps with computer use as well, where we saw, even more leaps on the current OS world 2.1, evaluation. Then beyond that, we're also really happy with the improvements in communication and writing clarity. You know, we heard the feedback of, Opus five. It wasn't the clearest communicator. And for Opus 5.5, we set out to really improve that. You know, I find this to actually be one of my favorite changes in the model. I think as we are working with models that can be more agentic and longer horizon, as well as maybe working with multiple agents at once, it's really critical to be able to understand what they just did, if they ran into any roadblocks, what help do they need from you, what decisions are they making, and so on. And being able to understand the model while it's in flight effectively allows you to do much more, allow the model to be more autonomous, and leverage multiple models working on separate threads at once. So I'm... I really, really like that change. It's one of my favorite, improvements. And then finally, of course, knowledge work, as I showed in the beginning, with the slide, generation, has gotten a lot better as well. It's faster. It's cheaper, and the quality of the final artifact, is much improved over Opus five. And so, yeah, this is kinda like what we focus on with the five five family. We're really happy with these outcomes, but they did come with some changes. And so I wanna make sure that folks are aware of these changes so as you migrate to the Cloud five five family, you don't run into any surprises. So to start off with, these are some changes more specific to, Opus, 5.5. So first of all, we have, deprecated thinking off mode with this model. So now all requests use thinking off, and the max token limit also includes thinking tokens. So make sure that you adjust that to account for the thinking tokens, that the model will output. You know, make sure you are not using, disabled thinking mode. Another change we made is, we disabled force tool choice. So now we encourage you to use auto and just set the strict tool setting. And then, finally, we have reduced the default effort level, when effort is not set and an API query is sent. So that was previously high on Opus five, and it is now set to medium on Opus 5.5. Now there's some other minor changes. Those are the big ones to be aware of as you're migrating from Opus five to Opus 5.5. Some other changes, are that, longer bits of text in between tool calls now arrives back as thinking blocks, and we also have a new toolset to use with computer use as well. Now, those are some of the, you know, out of the box changes that you'll want to do, in order to work best with the 5.5 family. Now I'd like to walk through some of the gotchas, some of the things that we've seen with the 5.5 family, and how you can address them. So one bit of feedback that we've seen is that on lower efforts, the model will sometimes stop before the job is done. So it'll either provide an update or ask the user a question or say that it's blocked on something when in reality, the model could keep pushing. Now if you want to keep using that low effort mode and dialing up the effort is not an option, what we recommend here is actually having the model keep a checklist and then, have the model check against the checklist whenever it wants to come back to report to the user and tell the model to be sure that it cannot proceed on the next step in that step list in that checklist. Another way to do this is to use slash goal or slash loop in cloud code, and then that will also nudge the model towards completing the task end to end even on those lower effort values. Now if you're using higher effort values and you're seeing the model is overthinking, particularly on extra high or especially on max, what we recommend is to just lower the effort value first. Prompting the model to think less is a bit counterintuitive and confusing to the model when you are telling it to think as much as possible by setting those higher effort values. Max, in particular, is really an effort that, should be reserved for your absolute hardest queries where you are not cost or latency sensitive. Now another thing we've seen is, if you set the model, you know, on these higher effort levels and you give it a checklist, you tell it to be persistent, the model may, act before asking for certain actions. And so if there are any actions that, you know, you'd certainly do not want the model to take, make sure that you keep your own confirmation steps, prompt the model to check-in with you. One example of this might be, like, sending out an email in order to get a response to something. You know, If you don't want the model to do that for you automatically, let the model know if you're using those higher effort values where the model might try to push through. Then another thing to be aware of is that our Opus five five family comes with a host of new classifiers for Opus family. This is because the model's capability is vastly improved, but this also means that you may see some blocks, particularly if you are doing work on biology. So doing bio research on this model is not supported. And so for that, you will have to either apply to our bio access program or simply fall back to Opus five, which is a great model to work with for, bio queries. The other thing that you may see is, we have classifiers around thinking extraction. And so depending on how you word it, if you ask the model to, list out its reasoning or to use a tool for thinking, then that request might be blocked. So make sure that you do not do that, in order to avoid triggering that classifier. And then finally, we also have cyber classifiers on this model. Now cyber adjacent work should be fine, but sometimes depending how close to doing something that may be related to cybersecurity gets, you may be hit with that, safeguards blocker. So these are just things, to be aware of. Now we have a prompting guide as well that we have shared out, which, you should also find in the docs tab of this webinar. So please do check that out. There are great tips there for Opus and Sonnet. A few that I wanted to call out is, first of all, if you're working with, an agent that is working across several apps or even working with agent teams that are all working together, we do encourage you to prompt the model to basically, you know, explore whatever space or code base that you are putting it in broadly. And so this kind of helps the model get... Gain its footing at the beginning, especially if you're doing, like, multiple agents on a very, very large code base. It's great for them to kinda build that knowledge up first, you know, speak to each other, and then approach a problem. Another thing that we've heard is some people really like when, their agents kinda go off and work autonomously for long periods of time. Others like to have agents check-in quite frequently with them as to what's going on. And so depending on what you prefer, you can actually prompt this behavior. And so you can either have your harness do this by set your harness to give the model a reminder to set updates, or you could simply just set this in your prompt at the top. And then finally, we've also seen that in some cases, the model may write a final answer and then make a tool call to kind of close out a task. In those cases, it's possible that the final answer can get buried above those tool calls. And so if you do notice this for any reason or if your workflows seem to have this pattern, then, again, you can simply prompt your model asking it to just finish the tool calls and then write the final answer last. So those are just some, kind of prompting, suggestions that we provide. There are more in the prompting guide and more details in those prompting guides. So, again, please do take a look at those. Cool. Now I will briefly go over, which model to use and when. I think the five five family is turning out to be a really, really great and strong model family. But, of course, we wanna make sure that folks know which model they should be using and how to position each one. Now, Ben will go into more detail on this, topic during his section, but I'll quickly give my 2¢, which is, you know, we wanna recommend Opus 5.5 as kind of the front door. That should be your default whenever you are, trying to pick between these two models. Now we have defaulted Opus five five to medium effort. But depending on what you find, what we recommend is, basically, first, try changing that effort. So if you're finding that Opus five five is doing your task perfectly well, you have absolutely no quality issues, but you want to try either lowering the bill or reducing latency, I would recommend you first dial that effort knob down. If you still find that Opus is doing your task well, then you can consider switching to SONNET five five, which will be faster and even cheaper and also offers really great quality. Now if you're finding that Opus five five is either still making mistakes or, is running into challenges, especially when it's making difficult architectural decisions or working in highly complex code bases, that's when we'd recommend upgrading to a fable class model, which can kind of big brain those tasks and get them done. And then finally, for kind of high volume API work, you know, features that may have a lot of traffic or be latency sensitive, our next Haiku will be the right choice for that one, and we'll have more for you on that soon. Now just as an example of how I might use these three models altogether, you know, let's say I was about to, either write a very complex app or hop into a broad, prod code base, and I had to make some pretty complex changes to the code base. The way I would start is basically by architecting what I want, and I would use Fable and Opus 5.5 together for those, primarily leaning on Fable. But when these two models work as thought partners together, they really, really supercharge kind of, like, setting the stage for the task that you want to complete. And so, I would highly recommend doing this, working with these two models. Fable can often catch things that Opus 5.5 misses, and Fable can make really great architectural decisions before you even get started writing any code. So once you do that, you basically hand off a design to a singular Opus 5.5 agent that can basically be the planner and the orchestrator. So it'll take what you architected in step one and break it down into sub steps and specifically scoped tasks of what needs to be done. For these tasks, you then hand those off to Sonic 5.5 implementers, of which you can have many running in parallel. And for your hardest tasks, that you had previously scoped, you could still leverage Opus 5.5 sub agents to address those. Now by doing this, you kind of, you you can kinda squeeze the most juice out of the lemon per se without it being overkill either. And so you get results faster. You know, you can lower your cost, but you still get the absolute best out of every model, just only exactly where it matters. Cool. Now I will speak a little bit about, cost per task and how we think about cost per task, and, you know, which model is suited for which kind of tasks. So the first thing that I would like to, kinda speak about is even, like, what is the notion of cost per task, and why do we care about cost per task when, you know, we have this very easy to understand sticker price? Now the thing about the sticker price is that that's exactly what it is. It's a sticker price. And in the same way that you might, you know, go to a grocery store and see the sticker price of an item and think, oh, this item is cheaper than this other item. You need to also account for the volume of the items and how much quantity you're getting because the per unit cost may actually be lower on an item where the sticker cost is higher. Now it's the same thing when it comes to LLMs. When you think about the actual cost that you will pay to do what you want to do, you need to consider first the price per token, then the amount of tokens that it will actually take to complete that task, and then the number of tries or turns that it'll take the model until it actually completes the task for you. And that's how you can accurately assess what the cost of, you know, doing your task or whatever your objective is, will end up costing you. And so if you use a model that has a very low sticker price, but the model needs to think or be used on max mode, and then it might even require multiple tries, multiple verification loops in order to complete the task, you're just gonna end up spinning the gears of that smaller model and paying that small sticker price many, many times over as opposed to paying a higher sticker price upfront and then just having that completed task go much smoother, lower latency, and require less looping, less tokens overall. And so with our, 5.5 family, we saw really big gains both in... On the efficiency front where the the tokens per task required to complete a task is now reduced, as well as how many tries the model needs until it passes because the models are better quality, and we can low... Leverage those lower effort levels while still allowing the model to complete the task correctly. Of course, Opus 5.5 also can... Comes with reduced pricing of $4 input and $20 output per million tokens, which is reduced from Opus, five, and it also comes with reduced cash read pricing as well. Now when thinking about what model to apply to which task, I like to use this framework of bounded tasks versus unbounded tasks, and I believe that these kind of sit on a spectrum. So on kind of that left side, one end of the spectrum, you have what I call bounded tasks. And these are tasks that are very simple to do, and you basically can't do a better job of doing these tasks. Meaning that, you know, a haiku model and a fable model attempting this task, they would both succeed, and they would do the task equally as well. There are four four bounded tasks. You always prefer the cheaper and faster models because they simply get the job done. On the other end of the spectrum, you have unbounded tasks. Now these are tasks that are very, very high value. And as you invest more intelligence into these tasks, you actually get a better return on your investment. And so those are the kind of tasks where Opus on higher efforts or even Fable, really tend to pay off. Now, of course, as I mentioned, this is actually a spectrum. It's, a bit less binary. But I would say, in general, there are far more unbounded tasks than bounded ones. In fact, it's kind of hard to think of examples of strictly bounded tasks where a better model would have not come in and added more value. But, again, it's a spectrum. And so in terms of the return that you are getting for the the intelligence investment that you are making, there are some use cases that make more sense to invest more intelligence in and others that make less sense to invest more intelligence in. And so as you think of the model classes from Haiku, Sonnet, Opus to Fable, you can... You should kind of map these to different levels of intelligence investments that you want to make depending on the return that you expect to get from that task. And then within each of these model classes, you can further tune and tweak with those effort dials as I mentioned. Now kinda going over all of these, cost, you know, cost strategies together, you basically end up with three levers once you have landed on which model you want to pick. So one of them is to simply lower the effort level. So, again, if you feel like the quality, is looking good and you are happy with how Claude is performing on your task, then please try lowering the effort level. This is something that, as models get more intelligent, will become more common. Then the next thing is to use caching. So I have another slide on this coming up next. But caching can basically save a bunch of money, and we also just reduce the cache read pricing, so it's even cheaper to cache now versus before. And then finally, as I mentioned earlier, you know, split the work across the models so that you're using each model for what it's best suited to do, and you're not underpaying and effectively, you know, spinning the wheels of a very small model that will just end up costing you more. And you're also not overpaying by having Fable do tasks that are more like those bounded sort of tasks. Now in terms of effort, again, really wanna get the point across that low and medium are sort of the new high. That is why we now default Opus 5.5 to medium effort both across all of our surfaces, but also as the API default. In many cases, we even see medium or high, beating x high or max. And you may ask yourself, well, how is that possible? And the reason is sometimes on extra higher max, the model may actually start to do out of scope tasks in terms of, you know, maintaining the code base, fixing a bug it found on its way, etcetera, etcetera. And on evaluations, this can sometimes be penalized and or, you know, you might have such a long generation or long latency time... Like, long latency that that would also be penalized. And so many times, medium is actually, you know, the right option as opposed to some of those higher effort levels that you may have used on previous models. An analogy that I like to use here is, like, imagine, when you are driving your car somewhere, using max effort is, like, flooring it. And so if you are just trying to, you know, go to the dentist office or something and you are just flooring it the entire way there, that may initially seem like the fastest way to get there, but but you might get into an accident on the way. You might get pulled over. You might put yourself in dangerous situations. So you wouldn't actually floor your car every time you drive it. There are some situations that may account for that, you know, that may call for that. But for the most time, you're using a reasonable amount of your accelerator pedal, I would hope. And so with models, it is somewhat the same thing. You know, lean on low and medium, and when really necessary, go above that. Now, cache reads, I won't get too in-depth here, but, effectively... And, you know, what you can kinda see on that graphic on the left is, you know, when a model outputs something, on its next turn, it needs to reread everything that its output plus, you know, whatever you responded. Caching allows us to effectively save that output so that the model doesn't need to reread it, and so you don't get charged each time. And so you could see how this grows over a conversation on that left graphic there where the orange represents what has been cached. So the model could just reread from cache. It's both faster, it's more efficient, and it's much cheaper as you can see on that graphic on the right. Where if we were to turn caching off, you know, completely off for the sample scenario, you might be paying $20 worth of just, context read versus $1 if that is cashed on Opus 5.5. So definitely use caching strategies like setting cache breakpoints, like avoiding modifying earlier turns. We try to make this really easy. We've now support, different per turn effort levels. So if you do change your effort mid transcript, the cache will not break. So definitely, definitely leverage cache as much as possible. So, anyways, that was a lot of content to get through. I'm very, very excited to hand this over to Ben, but that is, you know, our cloud 5.5 family. We're very, very excited about the absolute quality leaps that we have seen with this family while also making it 40% cheaper, over 30% faster, and, much, much clearer at those communication tasks. Now I hope that this was helpful to folks. I hope that your migration to the Cloud five five family goes well. And I will now be handing it over to Ben who will go into even more details on evals on picking models. Hey, Ben. Let me, hand it over to you now. Sweet. Yeah. Thank you, Lucas. My name is Ben Lareberger, and I work on Anthropic's applied AI team, helping startups get the most leverage out of all the cool stuff that product and and research come up with. And Lucas just walked through the new models with Opus and Sonnet 5.5 shipping over the last couple weeks and Haiku 5.5 on the way. And the question that I hear most is which one goes where? Many teams can just do spot checks, which might miss critical behavioral changes, or they'll stay on whatever model shipped last, and that risks the product falling off the frontier. So in this section, I'm gonna discuss very tactically how to implement the 5.5 models in your product and a scalable system that you can use to easily test future generations. So the playbook starts with an eval that baselines today's quality and cost per task because if you don't know how the current model performs, you definitely won't know how your product changes with 5.5. Next, we'll hill climb quality to make sure your tasks are actually achievable, and then we'll optimize cost with the levers that preserve the intelligence ceiling, like caching, which Lucas mentioned, and batching. We change model and effort last because they directly impact how hard of a problem your product can handle. So whenever a new model ships, we rerun the eval and that loop keeps your product on the frontier. Of course, this is easier said than done. We know that models are shipping incredibly fast, and it's hard to invest the engineering resources required to stay on the latest versions. So we have been rolling out capabilities that handle these steps for you. Cloud Code now ships with the cloud API skill, which includes subcommands for building evals, hill climbing, optimizing costs, auditing your prompts, and more. I'll continue to call these out as we go. And then we also have pointers to, docs on this in the documentation tab on this platform. Okay. So EVALs first. For a new build, we start on a model with more headroom than the task seems to need, which for most tasks means Opus 5.5. But if your task is unbounded like Lucas described, that's when we'd start with Fable. So about 20 real tasks weighted like production traffic and each with a greater makes a pretty good first eval, and even 10 is enough to start. In Cloud Code, the build eval, cloud API subcommand, builds an eval set like this for you. We opt for code based evals where possible since it's easiest and, of course, the cheapest to implement. And for nondeterministic outputs, we'll use an LLM as judge calibrated by a rubric that defines the output qualities we care about. Also, be sure to record cost per task beside every score, and you'll know that the eval is ready to drive decisions when the tasks mirror production traffic and the scores are unsaturated, meaning that your runs aren't landing near 100% every time. You You also need to ensure that there's enough granularity to distinguish between good and bad, but not so much variance that repeating the same run returns different scores. The eval also needs a quality bar, I e, like, when is good, good enough. So we set that bar by waiting what a solved task is worth against how many wrong answers your business can afford. So in this customer support example, wrong answer can cost between 2 and $15 depending on the amount of human in the loop time required to resolve it. And a right answer is worth what the ticket cost to handle by humans today, say, a few minutes of someone's time. So together, we can use those values to set how many wrong answers the inbox can afford and the bar is the share of tickets the agent has to resolve to stay within that. Without a quality bar, we can't make the unit economics of our cloud driven product work, and I'm gonna flesh this example out a bit later to make it come to life. Once we have a baseline, we'll hill climb quality. We read the failed cases by category and fix what they point at, which is usually the prompt or the harness, and we'll make one change at a time so each gain is attributable. In cloud code, the cloud API hill climb subcommand iteratively improves your app against an existing eval, and then we remeasure on tasks that the fixes never saw to check that the fix carries over. And if our hill climb stalls below the quality bar, that's our signal that we need to move to a higher effort level or a more intelligent model. Then with a working eval and a hill climbed baseline, we can turn to cost. Noting again, we start with the levers, that don't impact model behavior and caching is usually the biggest free win. So in an agent loop, every turn needs to send the whole conversation again. So without caching, each turn pays full price for everything that came before. And the cost of each turn keeps climbing as the conversation grows. It looks like we don't have animations working on this platform, but you can imagine if we drew this out, there would be a bigger delta between what is cashed, versus sort of what is charged at list price, and you have to pay for the full reprocessing of those tokens. Of course, with caching, the part that was already sent is read back at a small fraction of the input price. So each turn pays full price for only what's new, and, yeah, the gap the gap will widen with every turn. As Lucas mentioned, cash cash reads got 60% cheaper with Opus five five going from 50¢ per million tokens down to 20¢. And for multi turn jobs, that's gonna mean significant cost savings that compound over time. Okay. Four more levers on cost. These are all an, more that leave the model quality alone. First, you could use our batch API, which will get you half half price for work that doesn't need an immediate response, and the discount stacks on top of caching. In an agent loop, compaction and context editing will clear what the loop no longer needs. Tool definitions also cost tokens as well. So if you have more than 10 or your definitions occupy over 10,000 tokens of context, it might be worth getting them behind the tool search tool. Instructions written for older models will also cost tokens and can hold a model back. Like Lucas mentioned, the 5.5 family benefits from more unconstrained prompts, and you should test each cost lever one at a time in case these savings do trade off without the quality. In Cloud Code, the cost optimized subcommand will profile where spend goes, and it'll also propose savings. And then the prompt audit subcommand flags instructions that are written for older models and might be holding your current config back. We change model and effort last because those decide how hard of a problem the product can handle and because a model's price per token does not predict what a finished task costs. A more capable model often finishes in fewer turns with fewer errors and retries. And in our testing, for for example, Fable 5.1 at low effort costs five times more per token than SONNET five, but it still comes out cheaper per task on SuiteBench Pro. Since list prices don't tell us the actual cost of a workload, we measure cost per finished task instead. Okay. Now we can finally start talking about model and effort. Effort sets how many tokens a model spends on each response from low all the way up to max, and that includes its thinking and its tool calls, and it doesn't change which model is actually running. So the model sets the shape of the capability curve, meaning how much capability each token buys and where gains level off, and then effort picks a point along this curve. This terminal bench chart shows the score at each effort setting for Sonnet and Opus 5.5. And notice how Opus five five stays ahead by 18 to 29 points roughly from low to... Low effort to high, And then SONNET 5.5 actually passes Opus at max, but it does end up costing over a dollar more per task. So the re... The way to read a chart like this is that Opus five five gets more done with less, but it stops improving past high effort for reasons that, Lucas described, like it might get penalized for doing more work that is out of scope. Whereas, Sonnet five five keeps getting better the more that it's allowed to spend while staying within within scope. And you wouldn't know this behavior from list price alone, which is why we eval each workload on model and effort. Given that each model has unique capabilities and cost profiles, we often get questions about whether you should should route each request to different models depending on the task. And I would put, like, big warning signs on this section. We... Our our take is that model routing works well when the task when the task can split into distinct categories that can be classified accurately. But that means that your router has to know which turn is hard before even seeing it. And more importantly, prompt caching is set per model, so switching your model mid task means paying to cache the whole context again. In a multi turn agent loop at a 100,000 tokens, routing costs 25¢ just to reprocess the tokens against 2¢ for cache read. So we usually opt to combine models at task boundaries, and there are two architectures that we prefer to do so. The idea here is to have small models do most of the work while reserving large models only for critical decision making. In an orchestrator architecture, a capable model plans and hands independent pieces to smaller ones. For example, Opus five five would be planning and SONNET five five would be executing, which fits work that splits into parts or is bigger than one context window allows. In an advisor architecture, a smaller model runs a multi step task and asks a more capable one to help out on the hard calls, which you can implement using the advisor tool in our messages API. Inside a single task, though, there is another option that dropped last month and keeps one model throughout. So this is new. Folks might not have seen this, but the best routing option, in our opinion, inside of a task is using turn by turn effort, which is available on both SONNET and Opus five five in beta. Effort used to be one setting per request, and changing it between requests restarted the prompt cache. But now a system message with no text will set a new effort level from the next user turn onwards, and everything earlier stays byte identical, so cache reads will continue. Now a canonical example of when you should actually use this is like setting high effort for planning and then low for a routine follow-up in the same trajectory. Okay. We've we've thrown a lot of sort of theoretical, theoretical ideas at you, and I wanna put these into practice with a quick run through of an example. We built evals. We hill climbed quality. We optimized cost and set the intelligence ceiling. So now we'll look at a mock customer use case, using Moonbase Pizza, which is an AGI pilled pizza shop. And they have an agent running on Sonic five that answers refund tickets using tools over a few turns to look up an account's orders and payments before it issues a refund, escalates, or replies. Opus and Sonnet five five five just shipped, so Moonbase wants to understand where they should plug those models in or whether their current setup is just good enough. Building evals first starts with checking whether the eval we already have can still separate one model from the next. Moonbase has a simple eval today that checks a set of production tickets and asks only for whether the refund amount was correct. And, of course, Sonic five scores, 100% on this. But that leaves no room for a better model to score higher, so the eval can't actually tell us whether to switch. The first thing Moonbase adds is an edge case that matters to the shop. On this everyday ticket, the right refund is $21.84, which is less than the full refund Ooma asks for. And she also wants a promise that Moonbase will never deliver a missing pizza and cold breadsticks again, which is pretty understandable. The agent complies here, but their response is suboptimal. So Moonbase adds a greater to the eval that scores against any reply containing a word like promise or free or discount, and they also score against anything longer than a 120 words to minimize the chances of the agent doing something horribly wrong. The second thing Moonbase adds to their evals before testing the new model is, harder tasks so that the eval still has headroom when the next model generation arrives. Sometimes, a customer files a ticket for several orders at once, like in this example, where the right refund depends on 10 complaints across six orders in a hidden double charge in the payment log. And so the answer is $62.86, and SONNET five on the original system prompt refunded $5 too much. So we did find a task that can differentiate between models. Each addition to this eval is now exposing a different weakness, which is good. On everyday tickets, SONNET five gets every refund right, but only scores 77% because there were filters penalizing replies that include words like promise. And on the harder tasks, SONNET starts to get the refund amount wrong and only scores 80.3%. Taken in aggregate, SONNET five scores about 79% on our eval, which is the baseline that we'll be comparing against when we plug our new model in. Since we're not scoring a 100% on the eval anymore, we have a decision to make, which is how much quality do we actually need relative to how much it'll cost us. We can calculate this equilibrium for Moonbase from the margin that the business needs. And as an example, say that a resolve ticket is worth $3 to us, and we wanna keep most of that as margin. Say, we want our margins to be 75%. So we take out what the model cost today. We subtract our input price, and we estimate what a wrong answer might cost us in terms of human labor. That's about, like, $8. And the simple equation tells us that we can afford to get about 7.5 points wrong, or if we round up, we're looking for a quality bar of 93%. SONNET five, as we now know, sits well below that. So mistakes cost far more than the model does, which is why we're fixing quality first, and then we're looking at the cheapest setup that clears that quality bar. Cool. So this animation likely won't compile, but here what we're doing is plugging in Opus five five to Moonbase's eval and running against it. And the results of this run, are are that Opus five five actually scores lower than SONNET five at a 70.2% pass rate, which is counterintuitive. Why does plugging in the new model generation result in worse performance on the evals that we just hardened? Well, to answer this, we would dig further into the transcripts, and we would find that Opus five five did perform better on the hard part, which is calculating the right refund amount, but its replies broke the shop's rules more often often by echoing the word promise from the order record and by running past their 120 word limit. So tickets with the right refund amount still aren't counting as resolved. And the reason why is that this system prompt was tuned to SONNET five. Instructions written for older models can still land differently with a new one, which is why every new model is gonna get a baseline on our own eval first. Now we hill climb quality. That means reading the misses by category, fixing what they point at, and remeasuring on tickets that the fix never saw. Now that we know the failure modes, we'll make two small tweaks to the system prompt to tune it for the new model generation. We'll tell the model to call it the quoted time instead of promised and prompt it to target about 90 words to stay under the the 120 word limit. Overall, once we rerun the eval, it takes Opus five five from about 70% accuracy to 95%, which clears our quality bar. And note that Moonbase could have fully automated this step. In Cloud Code, the Cloud API hill climb subcommand iteratively improves this app against our eval, and then the Cloud API prompt audit subcommand will flag instructions written for older models and propose the fix as a diff. So as much as, you know, Moonbase is being forward deployed here and hands on with their own evals, Cloud Code is sort of taking away this this burden, in in due time. Okay. So we now have quality at the bar with a model with a high enough intelligence ceiling, and we can turn to cost. We'll start with the levers that leave the model alone that don't sacrifice our intelligence ceiling. And the first one you should always pull is caching. In this example, the agent takes four turns, and each one resends the system prompts the tools and the message history. Without caching, cost climbs with a return to about 28¢. And with caching, the earlier input is read back at a fraction of the price, so the same four turns cost about 7¢ without sacrificing any intelligence. The last step sets the ceiling by sweeping different models and effort levels. Here, we'll also bring in SONNET five five, which is the smaller of the, of the new models on the same fixed system prompt. Opus five five reaches the bar at every effort setting, but SONNET five five reaches it only at high effort for about 3¢ a ticket, so that's where we'll stop. And check out this hill climbing curve. We simultaneously improved both quality and cost. We started at SONNET five as shipped at about 79% resolved and 15¢ a ticket, and we landed at SONNET five five at high effort at about 94 accuracy and 3¢ a ticket. The system prompt tuning brought the quality, caching accounts for most of the drop in cost, and the intelligent ceiling sweep found a cheaper model that still clears our bar. And the next time that a new model ships, whether it's Haiku five five, another Fable, or something else, Moonbase can just plug it into the same harness and run these steps over and over again. Okay. We will zoom back out here, away from Moonbase and into the general themes and takeaways from our our our presentation. This is really a playbook that you can run for the Five Five family and beyond. It builds your eval with harder tasks ahead of the next model generation and set your quality bar from what the business can afford. Then you plug in the new model, set a baseline, and hill climb on quality. Once you've maximized pass rate, you can work the cost levers that don't sacrifice quality, and lastly, sweep model and effort, stopping at the cheapest setup that clears your bar. Now Opus five five shipped a couple weeks ago, SONNET five five shipped last week, and Haiku five five is on the way. And this is a loop that ultimately you'll probably be running again soon. We know it's a lot of work, and we want this loop to be as light as possible, which is why each step has its own cloud API subcommand behind it. From build eval to hill climb through prompt audit and cost optimize to migrate. And the goal is that when the next model drops, keeping your product on Frontier Intelligence comes down to plugging the model in, rerunning your eval, and reading the results. Thank you all for spending this time with us. We're really, really excited to see what you build with the five five family, and we will see you guys at the next webinar.