Video: Claude Code for Data Engineering | Duration: 3300s | Summary: Claude Code for Data Engineering | Chapters: Welcome and Introduction (4.965s), Cloud-Powered Data Engineering (211.89s), Integration Setup (376.385s), Production Demo Handoff (754.546s), Dashboard Investigation (793.4s), Fixing and Prevention (1473.745s), Sales Compensation System (1584.735s), Recap and Next Steps (2076.07s), Q&A Session (2189.52s)
Transcript for "Claude Code for Data Engineering":
Hello, everyone. My name is Kevin. I'm here from Anthropic working in data engineering. We're so excited to have you guys here. We're gonna share with you a little bit about how we think about Claude code for data engineering here at Anthropic and go through some of the things that we do here internally with the demo. So to kick us off, just some of the brief introduction. The data engineering team at Anthropic here, what we do is that we build Anthropic's data foundation and tooling, the pipelines, canonical data models, and platforms the whole company relies on. Some of our day to day work looks like the following. We work on data ingestion and integration, pipeline orchestration, data modeling and transformation, incident triage, remediation, code review standards, and data modeling and pipeline design. And for today's presentation, we'll be focusing on the right hand side, column of work that we do as data engineers. We're part of the data science and analytics organization here at Anthropic alongside the data science team whose core function is exploratory analysis, experimentations, modeling, forecasting, reporting, and others. And together, we help the company move faster and with more certainty through insights and data products. And now I'm going to quickly introduce the the presenters that you'll be hearing from today. So, again, my name is Kevin. I work on data engineering team here, and I'll pass it over to the other presenters for them to say a quick hello to you. Hillary? Hi. My name is Hillary. I'm gonna be doing demos a bit later, and I'll pass it off to Chandraprakash. Hi. This is Chandraprakash. I'm going to be coordinating the q and a at the end, and I'll pass it on to Chen Chang. Hey, Volks. I support data engineering and anthropic, and I'll be around, for some of the q and a. K. Awesome. So what we're trying to achieve here, today is that first, we're going to talk about the data engineering aspect of doing analytics with Claude. A few months ago, we published an article called how an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an an And so we're going to be diving deeper into the engineering aspect of that. Second is that we we wanna show you what a real day of doing data engineering looks like here at Anthropic, showing you how we think about the the the things we do in the morning, how do we fix pipelines, and how do we do design. We'll be using data dummy data, of course, but this this these are our real workflows. And lastly, we wanna share some actionable takeaways that you can bring back to your organization. Whether you're an IC or a manager, we wanna make sure that there's, something for everybody here in in the audience. K. And this is, again, the order the order of operations today, introduction, some technical grounding, demo, and closing. Some just brief housekeeping. The recording of the session will be distributed and available via email within twenty four hours of the conclusion of this presentation. Questions can be submitted on the q and a tab, in the webinar portal. And we would love all in any feedback you may have about this session, what you're interested in for future sessions. So please let us know how we could be more helpful. Alright. Without further ado, I would love to share with you about how we do data engineering the Claude way. As all data engineers know, fire drills can really easily take over your entire day, and there's a real cost to that. When there is a pipeline that breaks overnight, when there's a dashboard that goes stale, when there's a metric that moves mysteriously overnight, these can be really a big hit to our productivity. External research has shown that there are 60 to 67 data incidents per month for an average organization. Sometimes these innocent incidents can take nine to fifteen hours to resolve and sometimes even more time to detect. And as a result of that, so much of the data engineering's time are going towards data quality. And what we're proposing here is that data engineer can can offload a lot of these time intensive, reactive operational work to Claude code. That way, the data engineers can focus on proactive, high leverage work such as data modeling, pipeline design with Claude as a integrated and contextualized partner. Our hope is that a data looks like this, where you're just constantly finding fire drills, escalations, and operational task, where you don't have that time that you had planned to do data modeling pipeline design, can turn into a day that looks like this, where you can really have to dedicate a time and energy to do the data data modeling pipeline design, your plan actually happens, and your role in doing a reactive work would be more around verification and approval because Claude would be laying down the groundwork on that aspect already. So now we're gonna talk a little bit about how we can get there. As you know, Claude is really, really smart. We have some really, really powerful models here at Anthropic, and it's already smart enough actually to do both jobs. Anthropic can be really good at doing reactive work of tracing failure, opening the the fix for in in a PR. It's also really good at the proactive work of data modeling and pipeline design. However, that's not actually sufficient. Right? Well, what in order for for Claude to really maximize its impact for data engineering workflow, it actually needs to be integrated into how you you work within your systems and your workflows. Let me see the yeah. The okay. Sorry. There's issue with the animation here. And then it also needs to be contextualized in terms of following the norms and the processes of routines. In the next section, we'll talk a little bit about how we do that here at Anthropic. So first, we write integration. At Anthropic, we we build Claude into the workflows and systems that we already run. So what does that what does that look like? So the first piece of that is around MCP connection. So as a quick refresher, MCP connect connections are wrappers such so that a client such as Claude can interact with your your stack and your systems directly, would be it DBT, your warehouse, your orchestrator, your dashboarding tool, so on and so forth. And by having, Anthropic integrated to your existing systems, when something were to break, then the Anthropic agent can actually look at this problem from an ecosystem perspective, seeing how different different issues are connected together in a stage of PR for you to review. In addition to that, there will be a type tool called log, so that every all these actions are maintained and tracked. The second aspect to this is around access and trigger. What we do here is that we have set up, dedicated accounts so that a lot of these operations can be run with the right permissions set up. That way, we know that certain jobs with certain restrictions can be really enforced. And then after that's been done, then we can have triggers that set up at a particular scheduled cadence. This can be something that checks for a nightly nightly failure every night. This could be something that checks at a weekly or a monthly level. So and this can be very easily set up by Claude via a backslash Claude command or, other ways by setting up a cron job via via Claude. And lastly, at Anthropic, we like Claude work as how humans would like to work. I have colleagues who are really, really good with terminal. Myself, I'm a big Versus code user. Other people work interact with Claude different in different ways. So we really wanna make sure that our the data engineers on our team can really interact and work with Claude in whatever in whatever way that they're comfortable with. And that's a really big unlock for Claude is that you don't have to be migrating to a centralized way. You can interact with Claudes through different surfaces. And then after we have the technical layer set up, then we then we really have to give Kevin Luo the knowledge it needs to be successful as part of your organization. So what that looks like is that the procedures, the norms, and then the the the different definition of correctness become the files that claw is able to read. And the first piece of that is by creating the right type of skills for claw to do. A skill is a procedure codified as code for Claude to be able to run. So that skill could look like, you know, a morning briefing skill, which was which we'll be showing demo momentarily, or a PII audit or a pipeline triage type of skill. We talk a lot about this in a blog post, which is linked in the doc section here on the right hand side, where we talk about things like a procedural skill, a knowledge skill, so on and so forth. So please take a look at that. If you want more detail, we also have a template for a skill in the blog post as well. The second aspect around qualifying the norms and the procedures and correctness is early around the memory files. A memory file can be at an individual, a team, or a repo level. It can capture things like how you wanna do your SQL, how you want to write your documents, how you want to format your data models. But I think the really powerful thing here with memory files is that when you have incidents that have happened, the learnings that we have from particular data incidents can be captured in that memory file so that in the future when similar incidents were to happen, then Claude would know exactly how to react to that particular incident. Lastly, it's around semantic layer. This is, again, another topic that we fully, talked about in the blog post. We don't this this semantic layer is really a a a qualification of the metric correctness and the metric definition. So that way, when we talk about things like revenue, there isn't six competing definitions of what revenue is, and we can be really clear about what how we're defining revenue in what context and what's the supporting data sources for that particular definition. So that way, when we see a metric move overnight, for instance, we'll have full confidence that this is actually due to something that we wanna pay attention to rather than a shift in our definition. We don't really show this directly in our demo, but, again, our blog post will have a lot more information about how to think about a semantic layer. Lastly, I would just wanna mention that, you know, accuracy is really a context and verification problem, not really necessarily a code generation problem. And I know that is it could feel really overwhelming to set up the skills, the memory files, semantic layers in addition to the MCP, so on and so so forth. But you actually don't have to even write these files from scratch. Claude can be a great partner to you in drafting up the skills, the memory files, and then your role here is just really to provide guidance and to review. So after you have set up set set set up all these different things that we have talked about, a lot of the reactive work can actually be uploaded into Claude. Hillary Green-Lerman in her demo will illustrate what that looks like. And in that way, the time that you get back, you can be doing a lot of the design work that you want to do around data data foundation and tooling with Claude as an integrated and contextualized partner. So whether you're designing a pipeline, data model, a disoriented design doc, so on and so forth, Claude, because it's it's integrated, is able to check our current live production system to make sure that, okay, what we're designing is is not a problem that's already solved or that any changes has that has happened is it can be reflected in the the the current design. Secondly, it also has the context around, you know, how does the team like to, do the design review, how how does the team, like to do the documentation or any caveats or or knowledge. So then that way, you can you can you can go straight into having a well formatted document that meets the norms of your teams without without having to do that as an afterthought. And the last point I wanna mention is that these documents and these models that we design actually becomes tomorrow's context. And then that way, all future sessions can actually be will will know some of these work that you've already done, when when when they are look looking at, you know, the code base. Okay. So with that, I'll pass it over to Hillary Green-Lerman who will be, showing us all these different things, in production as the the other demo. Thank you so much. Hello, I think you're on mute. Sorry. Oh my gosh. I'm so sorry. I'm on mute. Hi, everybody. I wanted to give you all a live demo of my actual Claude code setup, but that's kind of hard given that all of the data I work with is proprietary. So I'm going to have to ask you all to suspend disbelief, as I walk through a mock up with fake data, URLs, and numbers. I promise that no real sales reporting or employee compensation data has been harmed in the making of this demo. So every morning while I drink my coffee, I run and review this morning briefing skill. I configured it to review orchestration alerts, chat messages, and even review live dashboards in the browser to catch potential errors. Let's see what Claude had to say for me this morning. Good morning, Hillary. Here's your morning briefing. Let's start with your pipelines. 29 out of 30 have fresh data. One stale pipeline roll up hasn't rebuilt in thirty eight hours. Oh my goodness. I didn't remember configuring it to be British. The most urgent item today is a dashboard mismatch. The overnight browser cross check found the exec and team pipeline dashboards disagree on q three numbers, 47,200,000 versus 46,800,000. This dashboard mismatch is likely caused by the stale pipeline roll up build that I mentioned earlier. We can investigate more when you're ready. Once you finish with the dashboard mismatch, there are two other urgent tasks. Finance team is expecting targets tables for today, and you need to merge p r hashtag one three two nine one as soon as the automated tests pass. Finally, there are 11 Slack posts to skim at your leisure, two FYI threads, a lunch poll, and two thank yous. Do enjoy those. Take a deep breath. We'll tackle the dashboard mismatch this morning and have plenty of time for deep work this afternoon. Let me know when you're ready to start. What you're seeing right now on the screen is a trimmed and anonymized version of my real morning briefing skill. This is something that I really did build with Claude and run every morning. The real one has specific pointers to Slack threads and BigQuery tables, but this gives you a good idea of what it looks like. And as you can see, it's just marked down with instructions to Claude, and pointers to different MCP servers that I've configured. The cool thing about a skill like this is that you don't need to write it yourself. You actually start by having a conversation with Claude about what you're working on and how you can make it more efficient. So the first thing you do is basically rant into your terminal about all of the repetitive tasks that you get do to get your day started. Which messages do you check-in Slack? Which orchestration alerts do you need to read and which ones are usually ignorable? Are there any quick queries you can run that tell you if a table is fresh and accurate? Claude takes all of these, all of this information and a few well configured MCPs and writes the first version of your skill. And then you run the skill and give Claude feedback on it. And it takes a couple of days of running this and giving the feedback and iterating on the skill, but eventually, it starts to get better. And I find that mine pairs really well with an iced oat milk latte. So the important part is after each briefing, you tell Claude if you got anything wrong and you add new data sources as they become useful. So for example, last week I started working on a new project and I added a bunch of new Slack channels that I wanted my morning briefing to monitor. The first few briefings were so so, but after about two weeks, I noticed that my productivity started to increase dramatically, and now I just can't live without the scale. So here are a few ideas of specific things to put in a morning briefing skill like this. From just a pure mental health perspective, the biggest win for me was having it triage my chat. I'm on East Coast time, go New Jersey, and I leave work at about 05:30PM to get my kids from day care. Most of my coworkers are on the West Coast, so they keep working for another three to four hours, which means messages from them start to pile up and pile pile up overnight. So by the time I log in in the morning, I usually have about 50 or more unread Slacks, and I start hyperventilating. This skill though knows which channels are the most important. It looks at direct messages to me as well as things that I'm tagged in, and it knows what projects I'm working on and what my general priorities are. So it's able to triage all of those Slacks and give me the top few, for me to work on first. It's also able sometimes to start triaging if the Slacks are, letting me know that something has broken. So rather than going in the order that I'm receiving these messages, I'm able to go in priority order, and I'm able to see that most of those 50 are actually not important, and it brings my blood pressure down a bit. Obviously, hooking up an orchestrator like Airflow, Daxter, whatever you use is super important, because then you can know which jobs have failed and Claude can actually start taking some actions there. It can clear tasks if there was just a transient issue, or it can do more more complicated workflows for you. I also give it a bunch of tables in my warehouse and some simple queries to run. Those work really nicely with the orchestrator alerts so that it's able to start diagnosing what might have broken, and it's also able to tell me that some of the alerts have already self healed. This is also really good for catching when there's an issue in the table even when all of the jobs ran correctly, maybe because there's some further upstream issue or a stream issue or a new failure method that we hadn't anticipated and written a test around. So this one's pretty self explanatory. You build in whatever alerting you have, like pager duty, and add those into your queue for the morning. Dashboards. This is one of my favorites, and I'm gonna demo it a little bit later. You can actually have Claude in Chrome pop open multiple dashboards, and you can tell it which numbers should tie out. So I use this if we have dashboards across multiple different platforms, maybe something in Salesforce versus something in our, internal dashboards. And this is a great sanity check because as we all know, even if all of the tails tables built properly, sometimes there's something last mile in the dashboards that needs your attention. And this is what's really catching things before an executive comes to me and says that the numbers are wrong. So code host, in this case, we use GitHub, lots of different, things out there. I, like to know what PRs are waiting for me. In particular, as we all know, when you start using Claude code a lot, your productivity goes up and you generate a lot more PRs and it you can sort of lose them in the shuffle. So this is a really good way of being able to track what has gotten reviews, what hasn't, what has merge errors from overnight. This one is very meta. You can actually have your Claude session look at other Claude code sessions that you are running. So if you have multiple Claudes for different projects that you're on and you feel kind of overwhelmed, like you're getting lost between all of them, your master morning briefing session can do an audit of those sessions and tell you where all of them are. You can also use this for handoffs. So one of the things I really like to do with the morning briefing is I'll get, sort of groupings of either failures or of tasks for the day, and I hand them off to the Claude that has the most context on them to start fixing. Everything that you've seen so far, and the fix that we're about to do is what I would call reactive work. Something broke, someone pinged me, and I'm responding. And that's most of my mornings. The thing Chen changed for me isn't the fire stop happening, although we've gotten a lot better and we'll talk about that going forward. But I found that the fires are now so much faster to put out, and I'm gonna show you what that looks like. So let's return to today's briefing. Chen had just told me that there was an issue with the dashboard and that it was likely caused by the pipeline roll up. Let's take a look at those dashboard mismatches and see what Chen saw in Chrome. So here's one of our Chrome tabs with a big top line number where we're seeing 47Point2. And then we go to the team dashboard, and we see that there's only 46.8 and that those two disagree. So I'm gonna ask Chen to start investigating the staleness in that pipeline roll up since that's where it thought the issue was coming from. Right. Pipeline roll up hasn't rebuilt because its upstream account eligibility failed every retry last night, and that failure isn't in your code. It's a Google Sheet. Oh my goodness. It's always the Google Sheet loads. It's the account review tracker. On Monday evening, someone typed the letters t b d in the target date column, perfectly reasonable in a spreadsheet. But that column is declared as a date, and since it's an external table, that one cell fails every red of the whole sheet. Row 5,415 column g. It's not the fault of whoever typed it, the sheet let them. So you need two things, one corrected cell, and I've put the link right there, and one task clear, so the pipeline rebuilds this morning instead of tonight. I've staged the command. It's a production right, so it needs your authorization. Okay. I'm going to go into that sheet that Chen sent me and see where the issue is. So if we scroll down, yep, we see that on Row 5415, we've got this TBD. Off camera, I've messaged the rep, and I know what the correct date is, so I'm gonna type it in. And now that it's cleared, I'm gonna ask Chen to do the fixes in airflow so that I don't even have to pop open the UI. Verified. The sheet reads cleanly again. I've cleared the task, and the pipeline is rebuilding now. The sync should be current within the hour. I've also drafted the update for your stakeholder channel. Shall I send it? Oh, that would be great. Thanks, Chen. Sent. Okay. So here's where my day gets interesting. The fire is out. What normally would have taken me hours to debug, we've done in a few minutes. With that extra time, what I would like to do is try to fix this class of bug and make sure that it doesn't happen again. So first, I'm going to ask Chen to act to go in and make sure that, these columns with dates can't get text in them anymore. So doing just some basic data via data validation in the Google Sheet. This is when we're making that turn from reactive into proactive. Done. That column now rejects anything that isn't a date with a note explaining why. And while I was in there, three other date columns on the same sheet had no validation either. I've added it to all four. No. I bet this isn't the only sheet in our repo that could have this problem. Maybe Chandraprakash can add that validation to all of the sheets in the entire repo. Yes. And there's some history worth seeing. There are 87 sheetback tables in the warehouse. 27 of them have typed columns a human could break with one keystroke. And this exact failure has happened seven times before on six different sheets. Every time the cell was fixed and the pipeline was restarted, the column was never hardened. There are zero proactive fixes in the repository's history. Okay. That seems like a really big problem, and we'd better fix it. Chandraprakash, can you have validation to the ones that we own and then reach out to the owners over Slack of the ones that we don't own? 18 sheets updated. Validation added on every risky column. Nine of them, I don't have edit access to, so I found the owners and drafted a message for each. Want to review one before I send them? Hi. Heads up on the budget tracker sheet you own. The target date column is typed as a date but has no validation, and it took down our org sync pipeline this week when, someone typed TBD. It's a two minute fix, and the link is here. Happy to do it if you'd rather grant edit access. Sent by Chandraprakash on behalf of Hillary? Great. Please send that. Sent. One more thought if you want it. Even with validation, a determined paste can slip through. I can change the pipeline's import to pass safely, so a bad cell becomes a warning row instead of a crash. That way the sheet can say TBD all it likes and nothing goes down. Shall I open the pull request? Oh, I didn't realize we could do that. Yeah. Let's definitely do it. Pull request opened. Amazing. Okay. So we fixed the problem for all of our existing sheets in our repo, but what happens when somebody uploads a new sheet? I'd like to be able to prevent this from happening We already have a code review bot that comments on each pull request. We can add missing date validation on sheets as one of the issues that it flags. That sounds awesome. Alright. Now we are using the magic of time travel to go into the future where someone has tried to load a new G sheet in that has a date column. As you can see, our review bot, which is running in GitHub with each PR, is able to catch this issue. I love our review bots. They catch things that I never would have thought of and they really made my code better. Okay. So our morning was all about fixing and preventing. The afternoon, we're going to move into designing. And this is actually my favorite part of the job, and it's the part that I never used to have as much time for as I wanted. So let me tell you a little bit about what we're going to work on. My main job at Anthropic is to provide data engineering support to our sales team. One of the most important tasks is coming up with sales rep compensation. Sales reps are compensated based on the performance of accounts in their book of business. The exact amount that they are compensated for each account is driven by a set of really complex rules, and those rules change every single quarter in order to incentivize different sales strategies. So this is one of the trickiest projects that I work on, and I'm gonna go through some of the reasons why it's particularly frustrating. So first is that every single number has to be explainable. It's not enough to just give the list of accounts that count towards someone's compensation. We have to also look at the accounts that didn't count and be able to explain to them why. Second, the tables need to run multiple times per day during the quarter and then stop running on the last day of the quarter before account assignments split for the next quarter. So every quarter, there's some changes to territories. A rep that was managing one account will manage a different account. And then also, as we said, the rules for compensation will update. So I need to figure out how to freeze at exactly the right time so that we capture all of the revenue from an account that should count towards reps' compensation, but not, so late that everything has already flipped over for the next quarter. Third issue is that there are tons of overrides. No matter how well we build our code, there are always weird issues where we need to override an upstream, datasets, ruling on who owns an account or what type of account it is. And we wanna make it so that I don't have to be the one entering in case statements every single time there's one of these overrides. We want it to be something that our compensation and panelists can do without code. So how do I go about building this? Obviously, I'm going to use Claude code. But rather than directly prompting Claude to start writing the code, I actually start with a design doc. And when we do that design doc, I start by giving a lot of context. I give all of last quarter's compensation code so it's gonna see how I like to write it. I give it Slack threads with all of the questions I've gotten from compensation analysts over the last quarter so that they can see what's confusing people. And I give the big requirements document that was written by the compensation team and is in plain English about how we're going to be compensating people this quarter. I'm also giving four problems that I wanna solve this quarter. The three issues that we talked about, one, I want this to automatically freeze so that I don't have to figure out exactly when to flip the PRs over. I want to make some sort of automated override system where people can just enter the overrides into a CSV without me having to edit any code. I want to make sure that the tables are structured so that rather than just showing the accounts that qualify for a particular compensation package, we show the entire universe of accounts and then have a human readable reason for each one about why it is or isn't included. And then finally, I wanna update the primary key for this quarter. Previously, our primary key was at the account level. I wanna change it to account and rep so that we can have multiple reps comped for the same account. I've read everything. All four are solvable, and three of them come down to the same idea. Make every decision a row you can see, explain an override. Let me walk through it. For the end of quarter freeze, you can use a project variable called compensation lock date. After that date, the models compile to select star from themselves. They literally reselect their own frozen output. For overrides, use one seed file per rule. New logo overrides, territory overrides, and so on. Each row has the account, the rep, the override value, the reason, and who approved it. The model reads that file as the very last step. The compensation analyst can add overrides without having to write any code. Here's the shape I'd propose. Capture every decision factor as a boolean. Did the rep own this account on the lock date? Yes or no. Did revenue cross the threshold? Yes or no. Then for every rep and account pair, we compile those booleans into one human readable sentence. This account counts because a, b, and c were all true, or this one doesn't because b was false. When someone asks why, the answer is already written. Would you like me to write this up as a design document with the proposed schema? Yes, please. That would be great. And I'm also gonna ask Chandraprakash to add a section on one of the issues that we had last quarter, which was how we decide if each account is eligible for the Chandraprakash code special payout. Added. Section 4.2 covers that eligibility rule with three signals, each stored as its own column and each contributing to the reason string. The design document is in the repository ready for your review. So here's the Google Doc that Claude generated for me, and I would then take this doc, edit it, and go back and forth with Claude until we had agreed on exactly how we want to build this system. This document becomes the contract that I have with Claude. And because I've spent time designing it really well, when Claude actually goes to build the code, it happens really, really quickly. So before I jump ahead, I wanna talk a little bit about best practices for building data models with Claude. So first, I try to state all of the rules for my model in plain English, and Claude will turn those rules into a model. Second, I describe edge case scenarios, and Claude writes them as tests. A lot of people don't think of writing tests for SQL code. It's not as common as when we're writing in a language like Python, and it can be a little bit more complicated. But that's where it's really wonderful that I just described the idea of the test, and Claude can actually do the mechanics of building it for me. So this year, when I was setting up our compensation, I gave a bunch of edge cases about different ways that, ownership could change mid quarter and how I would want compensation to handle that, and Claude was able to write all of those in. This is also good practice aside from Claude that it forces you to have conversations with your stakeholders and really nail down every single one of these edge cases to make sure that the code doesn't surprise you later on. Finally, when I'm writing the com the design doc, I try to nail down the schema with Claude at a really detailed level. I talk about if we have intermediate tables versus reporting tables. I talk about exactly which columns are going to be in each table. I talk about, like that example I was giving, what we want to include versus not include in each table. So now the fun part, I tell Claude to actually go and build the model, and it starts churning away. Now I would like us all to fast forward into the future, two weeks. We can imagine the code has been built and deployed, and now we have the the problem that I used to dread, which is somebody coming and asking questions about compensation. So I'm gonna read out a little Slack conversation here. We've got a rep named Kevin coming in and saying, hey. Why isn't Acme Corp counting towards my new q three new logo number that they signed in July? In this case, our analyst, who is not me and is not writing any code, says that he'll go check. He's able to go over into Claude, ask the question, why isn't Acme Corp counting towards Kevin's q three new logo number they signed in July? Acme's ISNI logo Boolean is false for two reasons. First, it crossed the revenue threshold last quarter, and second, it's not part of the top 500 accounts incentive program. Chandraprakash can then copy that back to the chat, share it with Kevin, and Kevin says, oh, that makes sense, and goes about his day. So thank you all so much, for going through this demo with me, and I'm gonna hand things back over to the real Kevin and Chandraprakash as opposed to our pretend demo folks, and we'll finish up. That was a great demo, Hillary. So today, we learned how to offload parts of your date, to Claude. We also saw a demo for both repetitive work and to scale yourself to fixations across proactively. Most important step to get started would be to configure an MCP and set up access. Keep in mind, without access, Claude's ability to work is extremely restricted. And every company has their own unique access setup, and, hence, this setup usually takes more effort than you expect. Context is how you tell Claude on how things are done in your environment. This is what makes it feel that Claude gets you. So we've we've shared, some takeaways for you. Try out what you learned today to automate one of your repetitive non fun work. Connect to an MCP, set up the access, codify the entire procedure that you do, and schedule it to run every day just as Hillary showed you earlier. You can also use Claude to brainstorm, some of your and to scale you in some of your proactive work just as, we showed you by fixing issues across the entire repository. Now we are ready to take any questions. Let me get to some of the questions that you asked in the, chat. So one of the questions which came was, how do you think about the mix of building and distributing reports, dashboards versus user drive insights via asking questions to Claude? Hillary, if you can take that question and Oh, I love that one. So I think of those as having two different purposes, obviously. We actually had this discussion internally about if we could imagine a world where there were no dashboards or report and everything was just in time back and forth with Claude. And we decided that there will always be a place for shared source of truth, and that's really what dashboards and reports are about. It's about everybody agreeing to look at KPIs in the same way on the same frequency. And it's also the easiest way to push alerts out to people, tell them about the data, tell them, what the frequency things are updated as. Now, obviously, the user driven insights, where you're sort of asking questions directly to Claude are also super, super important. There's a lot of different ways that we facilitate those. Having the semantic layer is super important to know that to make sure that what Claude says in the chat is aligned with those dashboard sources or truth. But I think they both have a place in the sort of AGI PIL post Claude data universe. Cool. Thanks, Hillary. Kevin, if I can request you on stage as well. The next question goes out to Chandraprakash. How do you ensure Claude delivers 100% reliable, accurate, and reproducible results? This is a wonderful question and appreciate everyone for uploading this. So 100%, reliability is a really high bar to achieve. And so there's a couple of things that we've, kinda highlighted throughout the course of this presentation. I I also wanna draw a plug to one of our blog posts that we wrote a little while back in case we haven't seen it about self serve analytics, in this AI driven era. But in short, a lot of the pieces that we've mentioned earlier today around skills, memory files, evals, and really truly doubling down on semantic layer, which we also just mentioned. These are all really key components to making sure, that high reliability and accuracy in your data, is guaranteed. Another push on the specific layer is making sure that that is well federated, well governed, and that there are single source of truth data assets that are being built. Those are all the bedrocks of a strong data foundation that will enable, you know, Chandraprakash at all to be, better decision makers. Thank you, Chen. Next question is, does having a super large Claude dot m d file degrade performance due to context overload? Yeah. So I can take this one. So we agree that context overload is a real thing. I think the guidance we have is that we usually wanna keep our MD files to about maybe, like, 200 lines or 250. I think that will be usually kind of the rule of some that we have. And I think that with every new model launch, it's actually a really good time to do a screen cleaning through your n d files. So what that looks like is you can act just work with Claude to say, hey. Can you see if there's any, like, stale, comments that are conflict conflicting or they're, like, no longer relevant to this version of the model that we can remove. So, basically, use the new model launch as a way to get rid of unnecessary context in these MD files would be one thing that we can do. Another thing that is you can do is that you can make sure that MD files are very, self contained in a domain and that can, like, link out to different other ND files. So imagine you can have one ND file for triaging or resolving particular data engineering issues, and then you can have sub sub MD files that link out to particular issues around check on delay or metric conflicts or anything like that. So then that way, the the the call, session only reads the MD file that it needs to resolve a particular problem at hand. Thank you, Kevin. Our next question is, have you reached the level of trust with Claude and this process flow where you are letting Claude remediate issues in the production pipelines, albeit with expected notifications messaging to a human. I can take that question. So as we saw in the demo, I'm very comfortable with having Claude do an investigation, come back to me, and request to take actions, and then I usually okay them. I think a big part of this is thinking through on my end what tests I need to be confident. So a lot of times with, you know, with remediation, if it's pushing a new PR to change things, that one's pretty easy because I can ask it to build things in dev. I can ask smart questions about the dev dataset that it built and then feel pretty confident letting it, you know, push the PR. For things where it's actually marking tests as passed versus failed, Again, it's about, do I have a good question that helps me feel really confident that, yes, this test was an error, and it's okay to mark it, as as as passed. But, yeah, as long as I have those good questions, I feel confident in it. Thank you, Hillary Green-Lerman. Our next question is, how long did it take for you to build all this environment around the implement of pipeline thing? How many engineers were involved? The kind of demo that you showed, Hillary Green-Lerman. I think yeah. So the environment so I'm assuming that that means, like, the background so that Claude could go and build the compensation pipelines. It was a very so there's a couple of of things that we can say when we mean environment. Some of it is just having a full repo with all of the SQL that everyone on our data science and data engineering team uses for sales data. That's not something that was built sort of intentionally for Claude as an environment. That's just the environment. But having that was really important because those are all of the examples that Claude has to work from. Then there are Claude, MD files in that environment that talk about our SQL standards, that talk about gotchas in our data. That was a sort of long organic process that everyone in our data engineering team worked on. So just to give some context, I've been on our data engineering team for, it'll be two years in October. When I when I started, I was the only one and I was part time, and our team has grown since then. And everybody sort of constantly is contributing to these markdown files, because our data landscape is constantly changing, and we're constantly telling Claude new things. And then so that's sort of the broad, can Claude write good SQL in our environment? And then the very specific, can Claude write good compensation code in my particular context? That was also sort of a process over time. The first time that I did it, it was a relatively light lift that I was asking Claude because I was switching over from sort of one quarter to the next, and the metrics were pretty simple. And and I was almost using Claude, as a very, very fancy rep to basically go through the code and switch up all of the dates. Then the next time I did it, I started to have these questions of, like, hey. There's this is a really labor intensive process for me to do the freezing. And I wish that I could have Claude come up with some ideas for how to do it. Similarly, we some of these metrics, we wanted to be able to keep multiple versions of it so that people could compare how it was calculated last quarter versus this quarter, and I bounced a bunch of ideas off of Claude. That worked really well, and it's kept expanding since then. And, Kevin, did you wanna add a little bit to that about testing evals? I think I'll actually probably pass that over to Chandraprakash to talk about testing emails. Okay. Yeah. I'll I'll take that as part of the next question. It says there's a question which says, can you talk about best practices when it comes to test evolves, and how do you go from 10 agents to 100 agents to feel comfortable with the changes being pushed without having to review every line of code? So, yeah, the you know, with Claude taking over so much, you need to understand where the boundaries is. So when it comes to tests and evolves it evolves, you want to implement deterministic test and evolves wherever possible. So, for example, in scenarios where you have CI checks, add as many tests as possible. There are certain, situations where you are using Claude in a repetitive, ways, like to answer certain questions, etcetera. And there are some aspects of those answers which can be deterministically evaluated, and you want to put the tests and the valves in place, there. But then there are all also scenarios where you cannot have those. And in those scenarios, how you then scale your 10 agents to 100 agents. And you have to build trust over time iteratively. So you start with something small with, Claude, and when it is able to do that for very consistently well. For example, Hilary did, her Excel sheet fix. And when she saw that it is able to do it fine in one place, she then validated it to do it at hundreds of places, etcetera. So that is how you scale over time gradually instead of doing it all at once. Keep in mind that there are certain kind of tasks this Claude can do very well and so there are others where it needs your help. And you build that intuition only by working with it again and again over time. Hope that answers your question, and it allows you to scale to multiple, agents. Let's go to the next question. There's a question on what is the difference between memory and skits? Maybe I I can answer that question, and perhaps Kevin can help, me after that. So the things skills are kind of your runbooks, things which you want of certain task to be solved in a certain way. It is usually shared across the entire team or at least, or even in the wider team or across the company. It's very well laid out on how you do certain things. Memory is something which can be only specific to you sometimes or sometimes very specific to your organization. So let me give you an example of memory for I want every time plot, is referring to a particular PR to be hyperlinked. I can say that next time, please put it in hyperlink. I want all my PRs to be in a certain way, to start in a certain way, meaning it should always start with the y first. And then I can put that into a memory, which may not be the preference across the entire team. And that's those are the kind of things I put into the memory. Sometimes the memory can have some things which you do specifically very differently in your company, which is handled differently in other companies, which can also go into memory files. Kevin, if you have anything to add. there. Sure. And then there's actually a follow-up question around how do we maintain skills, and co as a team. Should we store them in a repo? So I think the so how I think about memory and skills in the in the context of, like, maintaining as a team in in repo is that I think about skills as more procedural, step by step action. So, like, these are the things, for example, the morning briefing that Hillary has showcased. Now these are the things that we want Claude to do in order to produce that morning briefing, every morning for for Hillary. Right? And I think I think of memory as more almost like qualitative knowledge. So, like, let's say that we had an incident that looked this particular way. This these are the ways that we result a particular incident. You can ask Claude to document that in a memory. So in that way, you know, next time this were to happen, then it knows, like, it has seen this type of error before. And I think the value is absolutely having the story in repo and being committed. So then that way, you can see, like, how these, I mean, one way is that you can actually see how your memory has evolved over time. Right? You can see the different type of issues that you have resolved over the past, whether or not you're seeing these these issues happening again. So I think that's a really good way to leverage, the code base for that. And then another thing is really around colocation. So this is something that, again, plugging the the blog post, is that when we see our skills and our memory to be collocated with our semantic layer and then that we see, like, a significant improvement in the accuracy in our agent responding to analytics questions. So I think it's absolutely imperative for these to be, committed in storing code and and then share, amongst teams. But, of course, at the same time, you can have your individual, you know, skills or your individual memory file in terms of how you specifically like to work. But then I think from an organizational standpoint, it's really important to have that to be shared among team members. So that way, somebody else in the team is facing the issue that you have already faced before, then that's already stored in the team's memory file. Thanks, Kevin. I'll take this next question which talks about how good is Claude for troubleshooting pipelines, specifically spark pipelines, which can be huge? Like, the logs for spark can be huge. And, like, what kind of additional context does Claude need to be able to handle such kind of issues? Actually, Claude is extremely good at, troubleshooting pipelines, even Spark pipelines, identifying query issues, or performance optimization issues there. Now as long as you provided enough access so that it can read the spark log, it can correlate it with your GitHub. So it has the GitHub access where your code resides, and it is able to correlate them, on its own. You will see that it performed much better. Oftentimes, we see issues where people give it just the log access, but there is no access to the actual code which which caused that error in the log. And in those cases, it is gonna assume the code has something and perhaps will not give you the best results. So, hopefully, that helps. But the more access it has to be developed, just as the humans will need all of these kind of access to debug these scenarios, it will perform better. So our next question is where can we leverage Claude for dashboard development? I understand that we can do a lot of data engineering, but for analytics, especially on traditional tools like Power BI. Is anybody willing to take that? Hillary, you are on mute. speak. a little bit. I have previously used it, with with Loonker. I'm not super familiar with Power BI, but I can talk a little bit about using it with Loonker. Is that and I don't know if that'd be helpful. Do we think? I would say what and I think this is some general guidance too, that having really good underlying tables for the dashboards is a really big part, is a really big part of having good dashboards, and that's something that Claude is very well situated to. Also having a really good semantic layer is important for dashboards. And so if you're working with, one of the one of the tools, that has a semantic layer, Claude can be really good at writing that semantic layer and making it easier to build the dashboards. And then a lot of these dashboard tools also have some sort of text only representation, and Claude can write that as well. The other thing that I showed a little bit in my demo that I still think is, like, one of my favorite all time use cases is this is less about developing and more about maintaining, but having Claude in Chrome that can check between multiple dashboards and check if they're aligned is an absolute lifesaver. We have used that a lot, especially when we're going through migrations or when we have two different teams responsible for two different dashboards on different platforms. That has just been one of my all time favorite Claude and dashboard use cases. Okay. Cool. Looks like we are almost at time near time. I want kind of final notes. We're gonna post the recordings along with the slides and also a takeaway doc which you can use to try out what you've learned today. It'll help, like, your the starting points for for where you can start it. We'll also follow-up with emails with links just in case you've not been able to catch something during the, demo today. Thank you all for attending and for the great questions. Bye for now.