Transcript
Swaroop Chitlur: My name is Swaroop. This is Sidd. We started this team called GenAI Platform at DoorDash, and this is sharing our story about all the things that we did to make a successful platform. We talk about the principles, the bets, the pivots, and what do we mean by success. How many of you are shipping GenAI-based projects at work? How many of you are in a platform type of function where you are supporting other teams be successful at this? Hopefully some of the things we talk about will resonate with you all. Our journey actually starts in roughly April '23. As you know, November '22 was the ChatGPT moment. Then a few months later, we were like, we need to start using OpenAI. I was tasked with signing this contract, and I remember I was sweating, like, are we really going to spend this much money? Of course, in hindsight, that number seems minor now.
The feeling was very real because at that moment, I was like, if this thing takes off, how do we support this in such a large company? How do we make this productive? How do we make this accountable? That set the tone for what we were planning on. The thesis of this talk is very simple. We had initial principles and bets in mind. As the industry evolved and started doing newer things, we started adapting accordingly. We were still able to navigate these changes because we had those principles in place. Of course, we went from models to workflows to agents. That's where we'll walk you through our journey and what decisions we took, and talk about the technical projects in some of them. You get a balance of what was our strategy and what was the technical proof and what was the results. Hopefully those are takeaways for you at the end.
Operating Model
When we say about principles, what do we mean by that? When we started this team, we were given a blank slate. What do you do next? I don't know. Then you start thinking, first let's decide how are we going to operate? What are we going to decide on? We didn't come up with these. This was already part of our foundations or tenets. Be customer obsessed. Focus on the teams and the use cases. Not building for the sake of building product. Build products, not systems. Our customer is a product engineer. Don't make things where your customer has to stitch together different workflows. Try to think of an end-to-end workflow. Think of onboarding, think of support. Think like a product team, not just building systems. Make the right thing easy. If you say, do best practices, nobody's going to do it. You need to bake in best practices into how your platform operates. All this while we got to prove our value, our value to the customer and prove our value to our stakeholders.
Chapter 1 - The Customer Changed
Speaking of be customer obsessed, when we started with these principles, now we were like, we don't know what to build. We don't know what's coming. How about we talk to our customers and see what they're building and see where we fit into this picture. That's where we started. The first thing is the customer changed. I remember we were part of an ML platform team. The GenAI platform became a team under the ML platform org. We were used to thinking and talking to machine learning engineers. I remember in one of the conversations, I said, we want to try this out? Just use a notebook. They were like, what's a notebook? Then it just instantly clicked in our heads that we are no longer talking to ML engineers as our audience. Our audience is all engineers, which seems obvious in hindsight now. Remember, this is in '23, '24 that we're talking about.
That's when we realized, we need to move to APIs first. We need to think in terms of SDKs. We should not think in terms of notebooks or GPUs or getting access to compute. Thinking more in terms of APIs. Another decision we took in this process is that we'll focus on business impact. We don't want to build another chatbot. Of course, there were no coding agents back then. If you apply the same principle, we explicitly decided that this team was going to focus on business impact and product, not on the coding agent side, not on the chatbot side. That was a conscious decision we took. We wanted to double down on what kind of use cases our product engineers are building. As we spoke to them, we started seeing some trends, started seeing some broad buckets: automation, recommendations, personalization. Again, this was in 2024. We noticed that automation is decreasing the bottom line.
Recommendations and personalization is increasing the top line. We are in a general good sense of direction in terms of what value we can provide. We wanted to refine it further. Like, what is the value that we provide? That's when we decided that our USP was going to be help product teams optimize between these three things, accuracy, latency, and cost. You will see this theme emerge in all the projects that we took up after that point. What is the proof? Why am I saying that it was successful? Right now, we have over 5,000 users internally onboarded to our platform. Internally we have 45 users onboarding every single day. Surprisingly, something I did not predict was that 40% of people are non-engineers. Everyone from legal to sales and operations and strategy, they're all using APIs.
Chapter 2 - Adoption Created Accountability
Siddharth Kodwani: I'll go a little deeper into, from a technical point of view, what steps we took and what we learned from them. As in the last slide we mentioned that, eventually the audience that we are building for, we started with like, we built for ML platform, ML engineers. Then we went from, we are building for all engineers. Eventually we reached a point where 40% of the users of the platforms are non-engineers. It took us some time to realize that this new technology that is coming in is going to have a lot more adoption. It's not just engineers. It's going to have non-engineers, stakeholders are going to increase. We have to think from a platform point of view that we're enabling everyone. The ethos that we generally used to believe in while building such solutions is I want to build for velocity of my users. I want to build reliable solutions. As the surface area becomes big, you also need to have an accountability. I'll go in deeper for these scenarios.
As we started thinking about these scenarios, we need one surface that is usable with the whole company. The LLM Gateway is a concept that we came up with at that point. From an engineering point of view, we went to our users, understood what the problem was. This is early days. There are a lot of different providers. There is OpenAI models. Then came Claude models. Then came Gemini models. Then came open-source models. Now almost every product team is trying to experiment with them, they want to move fast. If they have to create their own plumbing for each of these different providers, each of these different types of models, then they were getting slow. From a velocity point of view, we quickly realized that we need one API and one SDK kind of approach where we do the plumbing for you. You get one API. You want to try an OpenAI model.
At that point, new model frequency was also pretty good. Every couple of months, you'll see one model that is supposed to be better than the other. Any team who is experimenting, they want to quickly try that model to see, is this application performing better with this new open-source model versus Anthropic's model versus an OpenAI's model. One API, one SDK, you make one change, and you can try a new model for you. At that point, it's important to be with the users. We were constantly with them to understand the problems that they are facing. We quickly realized that the capacity was still an issue. Even if you want to go use Claude model or an OpenAI model, you were getting rate limited. You needed good fallbacks. Teams are trying to build production services, they cannot have like, OpenAI has some shortage of quota or resources for us, and so we cannot use them.
We actually had to go understand what the fallback should be, what kind of routing should be. At one point, I remember we had OpenAI endpoint, but we also had a fallback on Azure endpoint. Or, we had Claude model API endpoint, then we had a fallback on Bedrock, so that the services can rely on us from the reliability point of view. We went in that direction. We quickly realized that as the usage is going up, then we need a way to track, because now finance got involved. Like, you guys have one API key for multiple teams, how do I track who is using what? We quickly realized we need workspaces where we have cost attribution. We can actually drill down on who is using how much. We can attribute cost back to those departments. That's what basically we did with LLM Gateway.
Swaroop Chitlur: Now that cost is back in fashion, we can talk about this. We had to take this into account from the beginning. Because if you're exposing this to everybody, enabling everybody, then somebody has to answer for who's paying for this.
Siddharth Kodwani: Talking about plumbing, LLM Gateway sat in between all the callers, all the teams who want to use the foundational models, the open-source models. This LLM Gateway actually boosted up the velocity of a lot of user teams, the users, services. We had accountability in place where any person could come and see how much is the usage for them, for their org, what's the budget for them, and what kind of model they are using, all that stuff. The principles on which we built LLM Gateway, the bets we took that we are building it for the entire company, not just for one set of users, actually helped a lot. Initially it helped with DoorDash. Then as DoorDash acquired multiple companies, Wolt, Deliveroo, SevenRooms, we were able to onboard them very easily and give the same kind of experience to all of them, so that they get the benefit of the velocity, the reliability and accountability principles that we have built in in our LLM Gateway.
Swaroop Chitlur: One of the positive side effects of this is that since we have this gateway, we are able to log every request and response. Now if an engineer wants to go back and see what's my performance, what's happening in production, then you have this UI that you go and look at all the logs. Every team does not have to build this on their own. The central platform provides this.
Siddharth Kodwani: As we said, make the right thing easy. We wanted all this from the point of view that we want to have a central control plane where we can have visibility, we can have attribution in place, we have forthright ownership in place, but we are still not compromising on the velocity of building. That turned out to be a pretty good decision that we made. The numbers speak for themselves. A lot of product teams are using LLMs in their products today with a very small onboarding time. They can revisit their tradeoffs. If you're using a different model, what is your latency? What is the cost associated with it? You can run those experiments very easily. You can debug those scenarios very easily. The adoption has been pretty good not just for DoorDash, but as we mentioned about Deliveroo, Wolt, and SevenRooms.
Chapter 3 - Vendor-First Taught us What to Own
Swaroop Chitlur: One of the bets was as an LLM Gateway, and that worked out very well. Another bet we had was going vendor-first. It's a bit counterintuitive because we are a platform team, we want to build, but we are also saying vendor-first. It was a conscious decision we took because this field is moving so fast. We could potentially say, we'll have this, but it'll take us six months to build. Of course, this was way before we had Claude Code and Codex. We took the vendor-first approach so that we can move fast, enable things, and essentially our velocity leads to the velocity of our product teams, our customers. It worked well for a very long time, and then it started to not work. Not all places, but there were a few specific failures that I want to talk about. Remember, we talked about accuracy, latency, and cost as our team's unique value.
How do you help teams increase accuracy? That's where evals come into the picture. Evals was controversial because every team has their preferred way of doing things. Nonetheless, we evaluated and onboarded a vendor. A couple of months down the lane, we realized that it was not working well because teams had very unique custom tracing needs, and the vendor was a good product, but it did not meet the flexibility requirements that we had. Team also had very custom annotation needs. Like now we have this image and we need someone to be able to annotate on this image what is it that they're looking for. These kinds of human annotation workflows is needed for labeling for us to know what is the ground truth, what is the actual result expected, but what is the actual value we are getting from the LLM. Teams started treating this vendor as a OTel trace store, and they were building their own workflows on top of it, which was least to say a suboptimal.
Then we realized like, ok, we are on the right track, but maybe what we have right now is not the right solution. In the end, we had to soften our stance on vendor first, and then we realized like, we need to own the workflow surface. What is it that our product engineers want to do? We own that surface, but how we provide that can be vendor or built in-house. We softened our approach to say, if the vendor works, scale it up. If it does not, take the lessons learned and then build a more suitable product in-house.
Similarly, on the model side, we went vendor first. Let's use OpenAI, Anthropic, Gemini models. Folks started getting productive. Folks learned how to do prompt engineering. They were able to scale up. At DoorDash scale, at some point, cost becomes an issue. There were a few other cracks that started to appear. Like the one is the deprecation treadmill. How many of you know that Gemini has a one-year lifetime for models? In one year, the model gets shut down, and this applies to every new model. At the end of that one year, you have to switch to a newer model, whether or not it meets your business needs. I could probably accept that if I was a single product team and I was working on my product, but I as a platform team, I have to think about hundreds of use cases and how I'm going to drive hundreds of teams to migrate to the latest model without necessarily giving them a business argument of the benefits.
Similarly, quota constraints. At our scale and at the rate of interest internally, we just cannot get enough bandwidth from these providers. Cost becomes a huge curve as the scale goes up. Go back to the GenAI gateway or the LLM Gateway, we had the portability principle in mind. You're able to easily switch models. Now it was no longer a nice-to-have, but it became a must-have.
Chapter 4 - Portability Became Real
It became a reality that, ok, what is the way to cut down? One of our long-term bets was that we can do fine-tuning and use open weights models and bring serving in-house again. We are an ML platform team. We love to build. That was our long-term bet. Then we saw that cost was becoming an issue. Then we jumped into our long-term bet, which was open weights models. Because we have the gateway, no significant change is required on the product engineer. They just change one config and use a new model. Of course, it's a whole lot of work to make that actually happen. We ended up prioritizing open weights models. We evaluated our vendors and choices. We finally ended up with Modal, which is a Python-first GPU cloud. We go through the LLM Gateway, and we are hosting our own models. The beauty of this infrastructure is that we want to bet on open-source engines like vLLM, SGLang for serving, and other open-source libraries, for fine-tuning like Hugging Face TRL, or Unsloth, or Axolotl.
We started investing this in the second half of last year. This year, adoption has gone through the roof once we had that foundation, because the models are getting really good now. Like if you say, Qwen, GLM, Kimi, DeepSeek, all of these models, they are getting really good out of the box. I would have not expected that as of last year. Now teams are able to switch from the frontier proprietary models to Qwen3, or, for example, Qwen3 Embeddings. They're actually seeing accuracy increase, and cost drops like 20x. Then teams are like, this is great. We should do more of this, because at our scale, cost becomes an issue. In the end, we were able to achieve cheaper and faster in just a handful of use cases onboarded in this first half. We already have single-digit million-dollar savings annualized. The vendor-first softened, but the portability held.
Chapter 5 - Agents Changed the Primitive Again
Siddharth Kodwani: Agents, how can we not talk about them? I think this is the time. The agents' journey also started in 2024, while we were building and solving a lot of user problems from the LLM Gateway point of view. We quickly realized that a lot more applications can be built with certain tool calls, with a certain harness, that can do a lot more than just LLM call. I'll talk a bit about it, like what was the context, where we went first, and then how we evolved from there. A little bit of context. As a platform team, sometimes, it's actually good to not be the first, like the flag bearer, and be like, this is what we are building from a platform point of view. Because we want to understand what your users are actually doing, especially when you're building in a GenAI space, where we have always worked with our product teams to understand what the blockers are, where can we come and help.
Sometimes you need to just take a step back, let them evolve and see where actually is the platform opportunities for you. Agents came in early 2024. Teams were just experimenting with it, building something. I think the critical moment that happened was when MCPs came, that's when a lot of people got involved here. They were like, we want to build this. Let's hook some MCPs and see what can we do from here. From there in 2025, we had 25-plus agents' projects, plus every laptop had an agent, like either Claude Code, Codex, Cursor, whatever. We started with, MCPs are here, we have third-party MCP servers, and we can start with building MCP servers internally. We understand all the learnings that we had from LLM Gateway. We'll focus on, again, the velocity of our users, reliability, accountability. We just focused on MCPs. Again, as it became very obvious that if you're building agents, it's not just about hooking up MCPs, because MCPs, it's like a presentation layer for the data to your agent.
If you're building agents, there's a lot more to it. What kind of experience do we want to unlock from agents' point of view? We started thinking from that angle. We basically understood that the users of our platforms are like, teams were building agents, teams were building MCP servers, users who want to use MCP servers or any other resources that we have on our gateway. Then there are agents running as services, agents that are running on behalf of users. A lot of different patterns came up. We quickly realized that we have to evolve from MCP gateway to agent gateway. I'll talk a bit about agent gateway. The agents are Claude, Codex, Cursor, plus now we have a lot more internal running agents. We started with MCP tools. Like, you can hook up Slack, GitHub, Jira internal MCP servers. Today we have more than 50 MCP servers running or onboarded on agent gateway. We took the lessons from LLM Gateway, like, let's build a very strong plumbing so that users can focus on building agents, building MCPs, and we figure out the way to expose them in a much more meaningful way.
Swaroop Chitlur: You don't want every MCP author having to think about authorization.
Siddharth Kodwani: I'll talk a bit about like, what were some of the core components that we started developing? We were talking with teams who are building MCP servers like, I want to authenticate users who are using my MCP servers, I need to have the right kind of visibility, because I might be exposing some data that should not be getting used by somebody who should not have access to it. We came up with authorization policies that makes it very easy for MCP servers to onboard on agent gateway. You can select the authorization policy, the tool level access that you want to enable for users. We will authenticate users for you. We will authenticate services for you who want to use your tools. We had right observability in place, not just for us, but whoever is onboarding to our platforms, so that they don't have to worry about the observability part.
We also enabled some of the features that were asked, like, my tools are sensitive for writes, and I want to have certain access level policies in place. We will rate limit your users. We built a very strong plumbing so that anybody who wants to expose any tool, they should feel comfortable coming to us. They should be able to see what tools already exist. We had a very good self-serve registry that you can go to, query, understand what already exists, what does not exist. If you are an agent, you can talk to us. We quickly realized that agents also want to talk to other agents. We were protocol agnostic where we are not only going to support MCPs, but we are going to enable agent experiences. We started thinking from that angle. We onboarded different protocols on agent gateway that you can utilize.
Swaroop Chitlur: Can you talk about AG-UI?
Siddharth Kodwani: To reach there, again, the principles that we learned from our past experiences, although the primitive changed, some of the things that we quickly realized that for LLM applications or the services, a lot of GenAI world is Python first. We advocated for Python internally in the company to have strong infra support, serve it as a first-class citizen, along with other language support that we had. That helped, because all the agent frameworks that came along that allowed our users to build agents quickly and start experimenting, they were Python first. We built products for our users, not just systems. If you're building an agent, if you're building an MCP server, don't worry about authentication, authorization, rate limits, integration with different surfaces, like Claude, ChatGPT, we took care of everything from the user's point of view, so that you can just focus on building agents and creating agent experiences.
Identity became something that you have to solve very securely from a safety point of view, from a security point of view. How do you enable an agent on behalf of you, and what permissions do you give to that agent to work on behalf of you? If you're working on your laptop with Claude, you can give a lot more permissions because you have a lot more control, but internally we have more agents that you can just invoke to trigger a certain action, but you don't want that agent to have all the access that you might have given to Claude Code. Solving agent identity tool access became a bigger problem that we solved for our users.
We took a controlled bet, the bet where we knew that this is not just about MCPs. We want to go in a direction where we want to enable agent experiences, not just give data tool access to our agents. We became protocol agnostic from day one. We focused a lot on MCP first, because that's where a lot of use cases were. Quickly, as teams were building, agents realized that they can take advantage of whatever the plumbing that we have built. They wanted to have different protocols supported. We were able to incorporate AG-UI. We are working on A2A. We expect a lot more agentic protocols in future that we can easily onboard and enable experiences for our users. Some of the results, because of the bets that we took, very quickly we were able to reach daily tool calls of 300,000. We were able to onboard new protocols like AG-UI, support streaming on UI directly via agent via the gateway that we have built. Python, the advocacy that we did allowed us to quickly enable our users to build agents. These were some of the results that were really surprising for us, but at the same time, very impressive.
Report Card
Swaroop Chitlur: We talked about the principles, some of the bets, and the pivots, and what was the success. These numbers are our success stories, along with all the use cases we enabled, like 5,000 users, savings, the amount of adoption that we have internally. Hopefully, these speak to what we mean by success. This is what we consider our own scorecard here. These were our bets. Product impact first bet held. A lot of the things that Sudeep's team shipped was built on top of the platform. APIs first bet held. This allowed folks to pick and choose the parts that they need. They are not locked into the platform, so we have much lesser complaints now. You can pick and choose the parts that you need. LLM Gateway was very successful in helping teams with their velocity. Portability was a bet that we had which worked out in the case of open weights models. Vendor-first, we had to soften a bit. We still feel it's the right first move, but we no longer consider it a must have. We do build when we feel it's necessary. New things emerged that we had not considered in the beginning, like evals, agents, and identity.
What's Next?
What's next for us going forward? This is how we look at it. From our top down, we still focused on agents running parts of the business. I know a lot of the talks have been about coding agents like Claude Code and Codex. We as a platform team, we want to enable things that run the business, and that requires a different mindset, a different focus. Maybe in a future QCon, we'll talk about how we enabled agentic commerce. We don't know. It's a bet. I don't know if it'll work out or not, or whether the world is ready for agentic commerce. We'll see. That's the top-down view. The bottom-up view is that we know what offerings we have. We have customers internally actively using it. They tell us like, there's this gap, or there's this bug, or can you add this functionality? We'll continue to expand on those. These are directional bets, but we keep watching for signals. What's happening in the industry? What product engineers are building? Where are they blocked? What are they telling us? What are they not telling us? We go from there.
Key Learnings
We hope to leave you with three takeaways or lessons learned. One, write your strategy early. It helps to have your focus, your bets. In a way, revisit your bets often, because then you know when something is working and when something is not working, then you consciously pivot and move on. The worst thing you can do now is make a decision and not reevaluate it. The second thing we would recommend is, build around customers. It's very easy to build these days. Everybody has their Claude Code and Codex ready on target. It's not always obvious what exactly is needed. All things remain the same. Talk to your customers. Third, own the company-fit surfaces. Auth is an evergreen story. Cost is an evergreen story. Accuracy is an evergreen story. Observability is an evergreen story. Governance, developer experience, and velocity. We own these things, but how we provide it is something we consider on a solution detail. Whether it's vendors or it's built in-house, those come later, but we own these surfaces.
Questions and Answers
Participant 1: I have a question on LLM Gateway. If you're still using, are you implementing the LLM safety and security in LLM Gateway itself? If so, how are you doing?
Siddharth Kodwani: When you say safety, security, are you talking about guardrails?
Participant 1: Yes, like prompt injection, output validation, hallucination control, and all that stuff.
Siddharth Kodwani: Yes. I don't think we can control hallucination at LLM Gateway, but we do have the concept of guardrails where you can come and define the guardrails. You can define the prompts that you should use, system prompts and stuff. From a safety point of view, the ball is in the court of the product teams to make sure that we are giving all the tools for you to go observe what is going on with the LLM calls that you are making. You can have a policy around it, like enable certain type of guardrails. LLM Gateway gives you the surface areas to go understand what is going on, debug it a little more. We don't enforce like, this is how it should be from a user point of view. That ownership is given to the product teams that are building solutions.
Swaroop Chitlur: To expand on that, the LLM Gateway is the right place to add before hooks and after hooks. Before hooks, for example, you want to do PII redaction. This is the right place to put that in a central place, and we run it for you. Post could be things like making sure that the LLM is not bad mouthing your customer or something, or recommending your competitor. That's where these were placed. For things like hallucination, we think of that more as an evals side of the equation. We have LLM-as-a-judge and deterministic judges there as the solution for that. In general, these two places is where we focus on guardrails. As of now, it's like customers opt-in and we don't enforce it. That's more of a company policy thing.
Participant 2: My question was around runtime or compute, any thought given to things like when you're trying to support it. One of the things we're struggling with is to try and maintain a consistent WebSocket connection between an agent session to a user, and it gets really tricky if your agents are running on Lambda or what have you, you can scale up and down. I'm wondering, was that like a consideration? What did you end up as your agent runtime?
Swaroop Chitlur: How do you handle streaming and scaling?
Participant 2: Yes. Streaming has implications on your agent runtime. If the first instance that you talk to, or the instance that answered the user's first question is not the same instance that captures the second one, then you can't keep a streaming connection. I'm wondering, was that a consideration, or have you faced that problem?
Siddharth Kodwani: This is like a developing area, at least in our company, like the agent gateway from the protocol point of view, we are stateless right now. We allow user teams who are building agents to treat us as stateless. If you have a certain tool that you want to talk to, they request a certain session ID, sort of handled by them. Overall, at the runtime, these are the challenges. As we mentioned, we work with product teams, we are seeing a pattern where we want to look at how the orchestration of agent can work, what's the runtime needed for that. We don't have all the answers yet. This is something that we are looking into now. We are talking to users, what are those scenarios where you want agents to run longer? What happens when the pod dies? How do we reestablish the connection? All of that stuff. That story is still not solved yet.
Participant 2: Session management, conversation history, ownership is still with the product team.
Swaroop Chitlur: Yes.
Participant 3: What were your key learnings or challenges when building identity and access management in the agent gateway? How traditional human identity or workload identity fell short when you were building agent identity?
Siddharth Kodwani: I'll focus on the learning part first. The learning part that we quickly realized is important, I think came from the fact that we are building a lot of new features for our users to quickly get access to tools. There are a lot of teams building agents, but they might not have all the security in mind from the user point of view. Let's say I'm a team who's building agents, and I want agent gateway to take care of whether this user can access this tool or not. It's like a three-party system where somebody has built an MCP server, it can have some write level access, it can have some certain confidential read level access. When such MCP servers get onboarded, we first tell them that, you should declare your authorization policy who have access to what. Let's say now that problem is solved. Now, a user wants to access something.
At the end of the day, what we are making easier for our agent building teams is you just tell me the identity of the user. I'll figure out whether this particular agentic action is possible on this particular tool or not. The learning was that, let's make it easier for agents to hand over that complexity of dealing with whether this user has access to this tool, and whatever the resource that tool might provide back to the user. The key learning was, let's make it easier for them. That's where we spent a lot of time on, how do we identify this agent? We started thinking in the direction where, we have an agent gateway control plane where a user can come and see that all these agents that are available that I use today, and these are the tools, so there is a cross combination where like this tool cannot access these particular tools for this MCP server, but this one can.
We enabled it and gave that power back to the user for them to also choose. There's two ways to think about it. MCP servers decide what authorization policy they want for the tools that they have. Then a user decides what agents have access to what tools that they already have access to. That lives in the control plane of agent gateway.
Participant 4: It's super cool that you guys have an agent gateway. I'm curious on how you think about the platformization of other aspects of the system, such as like, let's say the durable execution side of like, whether it's use any kind of workflow orchestrator and evolving them for LLM or agentic usage, as well as like from an eval perspective. I'm also curious that if 40% of the users are non-technical or non-engineers, how you think about the education process or introducing them to different concepts of when to use evals, when to use skills, and that aspect of it.
The first part is the platformization of all the other things that goes into agent building. Specifically, like whether it's eval suites, whether it's like CI/CD of evals, whether it's like the workflow orchestration, whether that's context storage, and thinking about all of that.
Swaroop Chitlur: For evals, we do have our own evals platform that is in the works. Nachiket on our team was leading that. Evals is another part of the story that we did not talk about today. That is the point where we focus on the accuracy aspect and we help teams build code-based deterministic judges, LLM-as-a-judge, and that whole loop of like product iteration. That part is platformized in our evals platform. With respect to the education part, I would say that still remains one of the hard parts. There are a couple of ways we focus on this. One is we do have a developer platform org that we collaborate and join forces with. For example, folks need to level up on how to access all of these MCP tools that we have. Or, you have, how do you figure out how to connect to Slack and things like that?
Our developer platform team has an agent skills marketplace. Most folks can just install that, and out of the box, their Claude Code or Codex or Cursor knows how to deal with all of the rest of the systems that we have. You can just install that plugin and just say, search for logs for this and it knows what to do. That's part of the education part. The other part is we have an active Slack channel that people ask questions and hopefully someday we'll have an agent answering there as well. Those are the combination of things that we look at. We invested a lot in onboarding and documentation. The truth is not everyone reads that. We attack it in different ways.
Siddharth Kodwani: We invested a good amount of time in coming up with the right skills because ultimately a lot of even the non-tech users are using all these agents, and if you have that right repo and right skills and right places, agent can quickly figure out what really needs to happen here. You don't have to understand everything. Your agent in most scenarios will figure out if you have the right skills available at the right places. That's where we're collaborating with the agent skills team. Where like, if agent control plane is the right way to call any tool today or agents, then how can you make it so easy for anybody to just do it? The idea is to just have the right skills in the right places and let the agents figure it out for you.
Swaroop Chitlur: In the end, we have to think about three audiences here. One is users, like humans, users on our platform, agents acting on behalf of human users, and agents as services. We think of these as three buckets of audience and we design for all three as we're building systems.
Participant 5: With reference to the open weight models, at what point in your cost journey did it start to make sense to get away from OpenAI or Anthropic and start to do some of the self-hosting and smaller models?
Swaroop Chitlur: For most teams, we always ask them to get started with the frontier proprietary models. One is they're easier to get started with. There's more information out there. There are actual people you can talk to if you don't know how to proceed. I think it's the three things: accuracy, latency, and cost. I would argue that those frontier commercial models are still the best when it comes to accuracy. It either boils down to latency or cost. If teams come to us and say, I have a very specific latency requirement, what do I do? Are products working great? We have product market fit, whether it's automation or whether it's external facing, but cost is a concern. I think that's when we drive the conversation towards, have you tried open weights models? Because you can shrink towards smaller models, you can think about things like distillation or other such fine-tuning space techniques to get to smaller models to achieve that latency that you need, or you can go towards a smaller model, for the cost aspect as well. Not everything needs the biggest model out there. Most teams don't realize that unless you walk them through the process. I think those are the two places where we usually encourage folks to start looking at open weights models. We don't encourage every team to start there unless they're familiar with the domain.
Participant 6: Can you elaborate a little bit more about your evals workflow? I see that you have a platform and this is great. From our experience, when it's easy to create evals, it's also super easy to get noise. It's really hard to get signal. How do you help your product teams to use evals the right way so they get signal and not noise? Specifically, I'm interested in CI, do you integrate it with CI? Do you have some frameworks?
Swaroop Chitlur: How do you get product teams helpful with evals? We're still in early stages. It's not a solved problem. Where we help teams get started, I think the first friction is tracing. We try to make that easy. Some of the work involved is like having agent templates. If you have the right tracing, like for example, if a team wants to get started, how do I build an agent? We have Google ADK. We have examples of how to hook tracing into the OTel store that we have. We first make tracing easy. Once that is unblocked, that's like half the battle. Because now you have those traces coming in, you can view them, you can understand where the time is being spent. People get a feel for like, how is the system actually operating? How's the agent actually operating? Is the bottleneck the LLM? Is the bottleneck the tool?
Is the bottleneck the results of the tool? Is it the web search is slow? Is it that the tool is slow? Once you are able to visualize what that observability tree looks like, they are more aware of what's happening. That's when you start questioning them like, what are the product metrics you are interested in? You cannot optimize on everything. You got to pick your battles. Like what are the two things or three things you care about as a product? Then, how do you operationalize that into actual evals or evaluators? We walk them through that journey. Again, this is a bit of handholding at this point, but hopefully over time we will learn enough to make this a playbook or self-serve. Once you are able to operationalize it to evaluator, sometimes you might need LLM-as-a-judge. Sometimes you can get away with just writing some code. You start building those evaluators, they are outputting metrics.
Now you need to visualize those metrics. You need dashboarding for that, custom dashboards. We have gone to the extent where product teams can vibe code their own UIs, and we will host those UIs for them on top of our platform. Including annotation workflows. Like if, for example, the agent makes a recommendation on these are the items we recommend for you at this restaurant. How do you know it's a good recommendation? Before you would have had Google Sheets where it dumps the entire menu on a Google Sheet, and you're like, you need a human to look at that cell and say whether this was a good recommendation or not. That's one of the bottlenecks we had, which is like, I can't visualize this. We enabled human annotation UIs, custom annotation UIs. This actually shows up as a proper menu. Like you can see the image, you can see the tag, and you have a custom UI with your custom questions on how to annotate this.
We have to take this step by step, get them started with tracing, get them started with what are your product metrics, set up the evals for it, set up the human annotations for it, dashboards for it. Once they are able to walk through this journey, it all clicks and you know how to do your product iterations. You're slowly hill climbing your way on accuracy.
See more presentations with transcripts