Video: Claude in Microsoft Foundry: Tool Integrations in Practice (APAC Rebroadcast) | Duration: 1808s | Summary: Claude in Microsoft Foundry: Tool Integrations in Practice (APAC Rebroadcast) | Chapters: Introduction and Welcome (38.904997s), Housekeeping and Overview (124.604997s), Cloud on Foundry (234.934997s), Five Core Tools (383.534997s), Foundry Cloud Deployment (476.034997s), Agent Demo Walkthrough (685.9249970000001s), Best Practices (1253.6849969999998s), Resources and Closing (1574.6949969999998s)
Transcript for "Claude in Microsoft Foundry: Tool Integrations in Practice (APAC Rebroadcast)":
Welcome, everyone. I'm Amanda Wong from Anthropic, and this is Claude in Microsoft Foundry tool integrations in practice. Here's the promise for what we're covering today. A good prompt only gets an agent so far. So to do real work, it has to reach your systems, read documents, pull in current information, and hand back something your downstream system can use. So today, we're going to build exactly that. One agent, three different layers of capabilities, all on Claude in Microsoft Foundry hosted on Azure. We'll also show you how to scope what the agent can do, what stays inside Azure, and how to get set up so your system is working the first time. Some quick introductions, I'm on the applied AI team at Anthropic. So I spend my time with customers building agents on Claude. So a lot of what we're covering today has to do with what we're seeing with customers building in production. And I'm joined by Haoran Chang from Microsoft who wrote the announcement blog post for the capabilities we're talking about today. So, Haoran, please feel free to introduce yourself. Yes. So I'm Haoran Chang. I'm a product manager on Microsoft, and I work on this collaboration between Anthropic and Microsoft to bring Claude model into Foundry platform and all the subsequent product coming along. So without further ado, I'll pass to Amanda to walk through the next part of the presentation. Yeah. So just some quick housekeeping notes. Don't be worried if you miss something from the webinar. We'll be sending out a recording within the next twenty four hours. We would also love your feedback, so we'll be posting a survey link near the end so that we can improve future sessions. And most importantly, if you have a question, please feel free to drop it into the q and a tab. We have technical experts here who will help answer those, for all of your burning questions. So without further ado, let's get into just talking about why teams are running Claude through Foundry. There are four reasons that we mostly hear from our customers, and the first is Frontier Intelligence. So this means getting access to Anthropic's latest Claude models, Opus, Sonnet, Haiku, which we'll also talk about in a moment. All of those are served with the same API surface and the same capabilities you would expect from Claude. So that includes coding, agentic work, and long form analysis. The second reason is built in safety and governance. That includes content safety, data handling controls, and compliance, which are all part of the platform. Third is integration with your Microsoft stack. So if you're already using Entra ID authentication, Azure RBAC, marketplace, and a path into m three sixty five or GitHub Copilot, the rest of the ecosystem pairs really nicely with Claude and Microsoft Foundry. And lastly, this is super important, but getting to pie from pilot to production more quickly. That consists of observability, evaluation, and all the operational tooling that goes into taking your product through the last mile. So we're gonna touch on these as we go, how we're gonna walk through getting a deployment set up in Foundry, and we'll spend time towards the end also talking about governance. But most of today is about the tools themselves, what you can build once you have a Claude endpoint in Foundry. So we are super excited to announce that Claude went generally available in Foundry in June. What going generally available solved was procurement and governance, Claude inside the same subscription, network perimeter, and cost management services as the rest of your Azure services. So there are two hosting options, and you get to pick which one you prefer when you create the deployment. The first is hosted on Azure. So inference runs on Azure infrastructure operated by Anthropic. Prompts and completions are all processed inside Azure, and there's a US data zone option that keeps inference specifically within The United States. Haiku, Sonnet, and Opus are all available here, and that's where most workloads should go. The second option is through hosted on Anthropic, where you get the full Claude platform feature set and the full model lineup, including Fable 5 for some of your most intensive workloads. This is all through your Azure account and built through Azure. So for today's demos, we are going to be using the hosted on Azure option. So let's talk more about the models available on Foundry so you can route each task to the right level of intelligence and manage cost without giving up the frontier. The first we're talking about is Opus 5, which is our frontier intelligence model. This is great for complex reasoning and long horizon agentic work, some of the hardest problems that need peak capability and autonomy. The second is Sonnet 5, which is our balanced workhorse. So this is where production workloads mostly start. That includes coding, multistep workflows, and customer facing agents. And lastly, Haiku, which we don't want to discount here in terms of capabilities, it's fast and cost efficient, great for high volume latency sensitive work, such as classification, routing, and extraction at scale. So all of these use the same API, same safety behavior, so you can prototype on one and route to the other without rewriting anything. For today's demos, we're gonna be using Opus 5 so we can showcase some of its best work. And in production, you may see a mix of Sonnet and Opus. So let's talk about the releases that, we are so excited to announce earlier this month. Five capabilities all on the Foundry endpoint and hosted on Azure deployments, as well. So here's what they are. Tool search picks the right tool from your library. Claude searches the catalog and only loads the few that it needs. The second is web search, which finds current sources and attaches citations to the exact spans that it drew from, and web fetch, which pulls the full document. So whether it's page text or even a PDF, it handles it just like an attachment. And we have web MCP connector, which calls your own systems. So that may be your CRM system, your ticketing, or internal APIs. By just pointing the messages API at your remote MCP servers, which you'll see in the demo. And lastly, structured outputs. This returns the expected schema back so you don't need to adjust it or deal with malformed output. All of these really matter because they can help you save time as you build your application and ship to production. Put this super well in his blog post. Highly recommend you talk through the reasoning and importance of some of these features. But the gist of it is that every team that ships a Claude powered feature usually spends extra time building this exact scaffolding. For example, a retry loop for bad JSON or a search and scrape service that has its own crawler and citation plumbing or even a hand rolled MCP client or tool router, that gets really difficult to choose the right tool past a few 100, tool options. So that's what the five tools are and why they matter. Before we start building with them, you need a Claude deployment in Microsoft Foundry and a first successful call, and Haoran is going to walk you through how to do just that. So thanks, Amanda. Before we get into the agent demo, I want to quickly show what the developer experience look like in Microsoft Microsoft Foundry, from choosing Claude in the model catalog to make your first API call in VS Code. So customer usually come to Microsoft Foundry because they want access to a front Frontier model, but they also need the enterprise capability to come with Azure. So the identity, governance, monitoring, and integration with their existing Microsoft main front. So what you're looking at here is the Foundry portal and you can actually see Claude Opus right here as a recommended model, but I'll walk walk you through the whole journey. Single to discover discover and go to model. That's our catalog. And you can go to source and choose Anthropic where you can see all the available model tiers. There's probably more. Right? So if you click, old models, you can see methods and and preview there since their data is not is not available in my project. But, anyway, let's click well, we're gonna use Claude Opus 5 days time. So let's click it in. Somewhere in here, you have some options to to play with. So first of all, there's two version of, and dropping models. So, again, Claude is available in Foundry with two hosting options, with hosted on Azure, the API and the model inference runs specifically in Azure infrastructure, while hosted on Anthropic, a customer can access to a broader Claude capability at this moment. For Azure hosted, we'll continue to add more features with the goal of reaching, parity with hosted on Azure Anthropic. And there's also different deployment option you can play with. So we'll just go with default, setting this time. It will ask you, will you use Claude four? Choose technology. Okay. So now what you're looking at is the Foundry Playground. On the left, you have, like you can play with system instructions, and you can add tools. You can add knowledge stuff, and you can also add memory to customize your agent in within Foundry. But for this time, we'll just do the very basic stuff and call the model. So we go here, and the there's a example code here. I'll just copy, and let me share my screen for VS Code. We'll just copy the code. And for this case, we'll use the Entra ID as authentication, which is more secure than API key. I'll click run. Alright. So Claude, the capital upfront is Paris. Great. I'll hand you I'll hand over to Amanda to go to the, agent demo. Thanks, Haoran. Now I'll show you how these come together in one agent. So here's our scenario. A field technician is standing on-site in front of a broken unit and wants to know three things. Is this part under warranty? What part do I need to replace it? And can it get fixed by a certain date depending on technician schedules? It's pretty simple question on the surface, but answering it requires checking a few internal systems, reading a document that might contradict the known part, confirming something on the outside sources, and handling a clean record so that the downstream system, whether that's Azure Azure SQL or Logic Apps, can schedule the work. So here's how we'll build it. One agent on Claude and Microsoft Foundry hosted on Azure in three layers of capabilities. First, the MCP connector and tool search so it can read our internal systems. Then web fetch and web search so it can read the manufacturer's notice and check outside sources. Finally, structured outputs, so what comes back is a record the downstream system can use. And now I'm going to transition to the demo recordings. In this demo, a Contoso field technician uses an agent to submit a work order for a failing industrial part at a Northwind Dallas location. Throughout this demo, we're going to use this exact same prompt to show you how the agent's ability to answer these questions improves as we add Claude capabilities. So we wanna understand, is this part still under warranty? How can it be replaced? And can it be fixed by Thursday? And if so, open the work order. Before we get into the details of the MCP Connector, let's take a look at what is available to the agent. So here we have a variety of tools available, such as getting information related to parts, such as when it was installed, where, and service notices related to that part. There are also other tools like getting stock, checking a technician's schedule, as well as creating a work order. Back in the MCP connector, I'll show you how we set this up. So here, all we did was pass in the name of the MCP server as well as its URL, which is currently hosted in an Azure Container app. That's truly all you need to do to get connected to the MCP server. Foundry connects to the server and runs the calls so that my Code never talks directly to the MCP. Normally, every tool's description, name, what it does, and arguments it takes goes into every request. But with this setting, defer loading set to true, call Claude only calls tools when it is needed. When Anthropic's Claude needs something, it searches the index, and Microsoft Foundry hands it that one description. Now note that the configuration requires at least one tool has defer loading always set to false, so I chose the one that is the most utilized tool, which is getting an assets details. So now let's kick off the prompt and see how Anthropic's Claude responds with the MCP connector and tool search enabled. So right now, Microsoft Foundry is connecting to the MCP server, pulling the tool list, and Anthropic's Claude is searching through those tools and calling them. Microsoft Foundry running a keyword search over the names and those descriptions of those tools that we shared earlier in the server setup. Microsoft Foundry then expands those references into full tool definitions in Anthropic's Claude context, so it's loaded only as needed. Let's take a look at the response back. So you can see that Anthropic's Claude was searching among the tools available to it as well as apply those actual tools through the MCP. And in the final response back, you can see that it noticed there was a warranty that covered parts and labor as well as the part that could be replaced with it. It also flagged that there's a PDF that corresponds to this particular asset, but it currently doesn't have a tool to open it. Now more on that later in the demo. And, also, notice similarly that it couldn't submit a work order on this part because there needs to be further verification. Now this answer is formatted, which is great for readability, but we might need something later on in order to submit work orders at scale to help structure this response. That will also be showed in a later part of the demo. And now let's add another layer of capabilities, web search and web fetch. And in this layer, you can see that we're using the exact same prompt and the same system prompt. The MCP server is set up, but what's new are these new two definitions at the bottom, web fetch and web search. Web fetch reads a document by its URL including PDFs. That's going to be really handy for the PDF that Claude noticed earlier, but it couldn't open. The second capability in this layer is web search. And you can see that it's fenced to a certain amount of allowed domains, which we've specified above, to be limited to the manufacturer's website, suppliers, and anything related to Contoso that we want to keep search limited to. Here's the document that we want Claude to read about that asset to determine what part needs replacing. And in this service notice that we've attached through the get asset tool and hosted on an Azure static website is that the original part actually needs to be replaced by another one with a completely different number. So web fetch needs to grab this information from where this PDF is hosted and use that in its final response back. Okay. So back in the terminal. Let's kick off the next layer of adding the web fetch and web search capabilities in case Claude isn't able to find the answer it needs in that document we shared earlier. So let's analyze the response back. In the first pass without web fetch, Claude assumes that we would need the same original part that ends in two twenty. But after reading the service notice through the PDF URL we gave it, it was able to replace it with the correct part. It was also able to use that exact part to check for inventory near the site. Anthropic was able to identify the right part by using the latest up to date information. Then it was able to create the work order and submit that. The last capability we're going to add is structured outputs, which is going to take this response from Anthropic on the right and turn it into a consistent format that our application can use. So I'm going to share this JSON schema that shares the expected output of how we want the Claude's responses to be formatted. It's a class that specifies the nine different fields, including the notes and citations related to web fetch and web search that we shared earlier. And we pass in the schema when we define the output configuration. So, really, what's all new from the different steps we took before is passing in this JSON schema. The schema complies to a specific grammar we're expecting, so the output can't be malformed. So now I'm going to kick off the final layer to apply this JSON schema. Again, we're using the same exact prompt, and we're going to notice how the response changes throughout and gives us back a structured response. Okay. So let's take a look at the response from Anthropic. You can see here that it structured it in the same schema we specified earlier, including the citations from the PDF that we were able to fetch earlier. The rest of the thought process such as searching for tools and utilizing them from the MCP server remains the same as expected. This JSON gets saved so that it can be passed downstream to the application with the correct part number and the rest of the fields that we wanted. So that's the agent. We added five capabilities, which really only took a dozen lines or so of configuration. Okay. So today, we focused on a field technician agent, but the pattern shows up across domains and industries. You can take this scenario and apply it, for example, to a practitioner asking about a medication or an adjuster asking about a claim or a banker asking about a loan file. The scenario changes, but the agent's job, the pattern capabilities don't. So now I'm going to hand it over to Haoran to talk more about governance. Thanks, Amanda. So, so far, you've seen the agent work of Prop three layer. First, it looks inside the compute the company using tool search and the MCP connector, then look outside using web search and web fetch, finally, I return a structured work order that another system could consume. Now I wanna shift from what agent can do to what we need to do to control before putting that agent in production. There are two important governance question here. So first, what made agent do? And second, what data leave Azure environment. So let's start with the first one. You can look at the left side of, the slide. So our field service, MCP server exposed eight tools, but this workflow only needs six. So the agent can retrieve asset history and existing work order. It can check technician and parts inventory. It can create a work order and assign the technician, but cannot delete an asset and it cannot update billing. So those two tools are disabled in the configuration, so they are not exposed to Claude in this workflow. And the principle here is pretty simple, We should not give an agent every capability just because the underlying system support it. We should expose only the tool required for this specific task. Alright. So the second question is with ALV Azure. In this demo, Claude is hosted on Azure, and our field service MCP server is also running in Azure, so most of this workflow stay with the Azure environment. The important thing is to watch when the agent connects to something outside Azure. For example, if the agent use web search or web fetch to access a public website, the the, the information needed for that specific request is sent outside Azure. The same idea applies to, MCP. Our MCP server is hosted in Azure, but if you connect Claude to MCP server hosted somewhere else, the tool request and response will cross Azure boundary. So that's why in production is very important to understand where each tool is hosted, where data what data it receives, and whether that create an external data path. And for public web access, you can also restrict which domain the agent is allowed to reach. Anyway, the main takeaway here is know where know where your tool is are running and know when your data cross Azure boundaries. Last but not least, there are six things I want to call on to get things right for the first time when you're working with, Microsoft Foundry. So first, first first thing, send the version header. So if you call the message API directly, please include the Anthropic version header. But if you're working with SDK, that normally handle that for you, but it is easy to miss with a raw rest request. And same thing goes to when you are trying to use the tool we talked about. Most of them require a sort of version header or beta header. Second is set max token, explicitly. This is pretty self explanatory. Max token means the maximum number of output token Claude and generate. So if the value is too low, the model can stop before a meaningful response is complete. And application should definitely check the stop reason and handle MaxToken properly. So third thing is defer tools and add tool search. So Claude does not need every tool definition to load it into its context at the beginning of every request. So smart smart way to do it is to keep the tool search tool available upfront and mark the larger tools as deferred. Now a Claude can then search and load the relevant tools when needed. One important rule here is that tool search itself must remain available. Fourth one, allow allow us your MCP tools. So we wanna start from default deny and explicitly enable only the tool required by your workflow. You know, demo, that means enable the six field service actions while excluding delete asset and updating update billing. Okay. Fifth, keep your output schema straight. For the final work order, we use output config with a JSON schema. So you have to define the required fields really clearly, use the correct data types, and set additional properties to false for each object. And last but not least, I wanna call the schema cache. So structured outputs works by compiling the schema into a grammar. So the first request with a new schema may take a little bit longer while the compilation happens, and the compiled grammar is then cached, making later requests faster. After that, it's cached for, like, twenty four hours. Please version your schema so it change, doesn't silent silently re reuse a stale one. Okay. That's it. So Thanks, Haoran. So that's everything we plan to show you today. We wanted to share a few resources with you so that you can go and get started and deploy Claude in your own Azure account. So a couple resources are the GA announcement on the Microsoft blog to learn more about how to use Claude today. And just as a side note, if you're interested in these slides, we will share them with the recording. And you can also access them now in the resources tab in as you're viewing this. Haoran will also talk about the other documents we have for you. Yes. So to view Claude as soon as possible, you can also go check out the blog and write about the five tools that we just recently launched. It's on Microsoft dev blog. Also, you can just scan the QR code and see the Claude tool docs. And that's it. Thank you so much. So much for joining us.