BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage Presentations Beyond Observability: Evolving Production Operations in the Age of AI

Beyond Observability: Evolving Production Operations in the Age of AI

01:01:17

Summary

The panelists discuss how production operations are evolving with AI, turning operational data into actionable insights for incident response. They explain how automation and new architectural practices help engineering teams build more understandable systems, while exploring how AI reshapes software delivery and changes the historically deterministic nature of production applications.

Bio

Michael Hausenblas is a Solution Engineering Lead in the AWS open source observability service team. He covers Prometheus, Grafana, and OpenTelemetry upstream and in managed services. Sujana Sooreddy is a Software Engineer at Netflix. Noam Levi is the Field CTO at Groundcover and one of its founding engineers. Renato Losio is InfoQ Staff Editor, Cloud Expert and AWS Data Hero.

About the conference

InfoQ Live is a virtual event designed for you, the modern software practitioner. Learn from extraordinary speakers driving change and innovation.

Transcript

Renato Losio: In this session, we are going to chat about, "Beyond Observability: Evolving Production Operations in the Age of AI". In this roundtable, we're going to discuss how production operations are evolving, how it's been actually changing in the last couple of years, and how AI and automation can help turn data into insights. Basically, what practices will help teams build systems that are easier to understand and operate. Our panelists will explore how AI reshapes our work with production and software delivery and how it changed the deterministic nature of our application.

My name is Renato Losio. I'm an editor at InfoQ. I'm joined today by three experts, coming from different companies, different sectors, different backgrounds. I'd like to give each one of them the chance to introduce themselves and share their professional journey in observability.

Michael Hausenblas: My name is Michael. I'm a principal software engineer in the SRE team, part of platform engineering here at Genesys. Before Genesys, I worked seven years in AWS, various service teams there. That's where my observability journey really started, specifically around OpenTelemetry.

Sujana Sooreddy: I'm Sujana Sooreddy. I'm an engineering manager at Netflix, working on media systems and observability. We are part of the live and encoding technology, and we provide the infrastructure and observability that let anything you're seeing on Netflix go through these workflows and make them efficient, and have the infra visibility in it. Lately, a big part of that job is just making our systems and workflows AI native, both using AI to run our own operations and making sure our platforms are built so for the agents. That's what we're up to these days.

Noam Levi: My name is Noam. I'm one of the first two founding engineers in groundcover. These days I'm the field CTO. groundcover is an observability platform, which I think these days can safely also be named an agentic observability platform with the focus on privacy via a BYOC architecture and eBPF sensor, which I'm sure that Sujana from Netflix also has some knowledge about eBPF as Netflix are one of the biggest contributors to this technology.

Role of AI in Observability

Renato Losio: What's your experience, and where are you currently using AI in observability today? What are the practical problems that you try to solve first with AI in observability?

Michael Hausenblas: The journey really starts at the development side where through skills, things like, for example, OpenTelemetry instrumentation, manual instrumentation can be supported. That goes all the way into troubleshooting production incidents, or after the fact really figuring out how can we improve the situation overall. We see that across the entire lifecycle really. I think the overall gist or the thought that I want to implant here from the get-go is really this trust but verify. It's an old Russian saying, like it's good to trust, but always check and verify. The how is the tricky part. I guess everyone can agree, but how to exactly do it, I think we'll explore that a little bit in greater detail.

Sujana Sooreddy: For us, there's like the individual productivity that comes through writing the code, building systems while developing. The biggest wins that we are seeing are more on the outer dev loop and the operations aspect of it. The clearest example with the big wins is like rolling agents out as first responders in all of our support and alerts channels. This is basically not just the support of platform, but also application and product teams as well. That is where we have seen a great reduction in the amount of time, that is the MTTR to resolve the incidents. Just figuring out and answering the support questions, that is where we have seen the institutional productivity gains, in that way. In terms of the development loops where we are, one thing, fascinatingly what we're observing is like different models. Whenever you change models, whenever you move, it provides very different style of coding. We are working through how do we provide the context evals and harnesses in terms of the inner dev loops. Deployments, alerts is where we are seeing the biggest wins so far.

Noam Levi: As far as I perceive, it is a three realms issue. The first realm is the production issue, where we're trying to incorporate observability to lower MTTRs and to resolve support tickets as fast as possible. I think where we also see a very interesting motion with AI at least internally is during the development lifecycle itself, where I would say that it is safe to assume that today's development lifecycle velocity is unprovisionable with AI. Meaning that code is moving faster than the engineers that ship it can internalize it. This forces us to also adapt the model where observability is part of the development lifecycle itself from the initial step where the agent starts to print its first output tokens in the form of code. That will be the second motion. I would also say that the last motion that we see, which is interesting, is that observability data becomes more and more relevant in many parts of the organization where it used not to be that relevant in the past, or was harder to incorporate into this part in the organization. We see at least internally in groundcover where even non-engineers are asking, in our case, the groundcover platform, but observability data sources in general, questions about the application, which are directly relevant to business level questions in other parts of the organization.

AI Use Cases, in Incidents and Alerting

Renato Losio: Actually, as a user, I'm a cloud architect playing mostly with AWS technology. I've been sold the idea, I'm an end user in both cloud services, non-cloud services, third-party. I've seen three major areas that have been targeted as a potential advantage of using it. One was alert overload. One is you're going to be faster investigating your P1 or any incident at 2 a.m. The other one is whatever incident you're able to summarize and get a full report or whatever else. I was wondering which one you find easier, or which one you think is already there and I should fully take advantage of it. What is really just still something we dream about? Because I get it that, yes, someone can filter my alert. I can see the benefit, but I haven't really been able till now to bring it to production myself.

Sujana Sooreddy: If you are focusing on incidents and alerts, the biggest thing that you could do is, previously so far in all of our distributed world, all of our alerts and observability systems are written for humans with like graphs where it fancily loads and catches the eye. If you want it to be successful with the machines, it's the raw data. The more and more you could come up with the raw data with easier to read. I think naming is becoming even a bigger challenge, which is already a challenge all the time. Those right names, right context. For it to actually have those enough, like the harnesses and evals that are present so it can stitch between different tools. I'm pretty sure everyone as an engineer uses half a dozen tools as part of triaging, going from like an alert to logs to traces to metrics. How can an agent navigate through all of this journey?

It becomes more important. Enabling that agent to do that switching from metrics to logs, when to go from where to what, really helps to get those productivity gains in terms of alert debugging. Incident retros, I think it's so great. It's because it's more inference. It's less on actually looking at data. For incident retros, you could also make a skill which makes it look at being a non-blame approach. How should you look at it? How should you track through those things? It works really well there as well.

Michael Hausenblas: There are two aspects I'd like to highlight. One is really around risk and the other one is around accountability. With risk, we, for example, in Genesys providing this customer experience operate with highly confidential data. There's a bunch of regulation. There might be things where we actually need to provide auditors certain answers, provide evidence that someone didn't have access to whatever data. There, the risk really is leakage. Like if you have any kind of PII data that might leak some PII, personally identifiable information, think of your address or your email address or whatever that might leak. That leads me to the accountability part. If I have a human operator in the loop, I can essentially go to that person. It's like, this is Renato's fault. Renato leaked it. He fat fingered it or whatever. Consequences aside, but I have a human essentially there. In a low-risk environment, if I look at that test.

Of course, not a problem or not that high of a risk. If you think about these kinds of potential leakages, at the end of the day, you probably want to have a human being accountable for that. That doesn't mean that I need to be alerted at 3 a.m., and jump in front of the laptop. If it's a simple thing like, ok, seems there was some issue, I can roll back to the previous version. That is something that at least we need to pay a big part of the attention to and do on a daily basis.

Noam Levi: I want to give maybe not a different perspective, but put maybe things in a larger perspective. My feelings when it comes to AI use cases are mostly that we are currently in the T0 of AI in general, which means that I think that changes are too frequent to be decisive on whether AI can be more practical in one part rather than the other. Like if I will take an example from the frontier model labs, a few months ago, we were all using probably like Claude and Mythos was considered to be the top LLM to use. Now there is all this chatter about Grok Bot and Muse and Astra. Everybody adopted MCPs and all of a sudden CLIs were all the rage and MCPs were deemed obsolete. All of a sudden MCPs are rebranded as connectors, and everybody adopted them as well. I think, not necessarily AI, but as we are reaching what people like to call AGI, reaching a stage where AI can be really productive in any part of the business, I don't think it is a question of where AI is more productive.

It's then how hard we push teams to generate signals that AI can work with so they can enjoy AI. I think that traditionally SREs or at least DevOps teams are just early adopters of technology and have a better ability to work with systems that still show friction. We are currently in a place where the friction went below a critical threshold, where I would say that AI can transform finance operations and middle management issues as much as it can solve alerts in production. It's just a matter of making sure that the right connectors are in. The culture is about generating data and serialize the company's work into knowledge bases that AI can work with.

Shifting Operational Complexity, with AI

Renato Losio: Do you see AI at the end reducing, so the operational complexity for me as an SRE is actually moving somewhere else?

Noam Levi: Yes. I think that the complexity seems to be more and more, at least in teams that we see, let's call them AI-first companies, companies that are built into this era where AI is building the code from the very beginning. I can tell you that as an observability vendor, those bleeding edge companies are reporting to us that their interaction, for example, with the platform has almost become fully, if not completely agentic. Some companies report more than 80% of the adoption of the platform was agentic. At this point, their job becomes more around managing their own human context when they're switching through different agent runs, making sure that the agents have enough context to complete the job, as Michael said, being accountable for the work that these agents are doing. Accountability. This is the main focal point, at least of the pressure that these teams are handling. I think it's mostly around that management of context switching and the ability to create safe sandboxes for the agents to make decisions and make sure they have enough context to take the right decision.

Good Production Engineering in the Age of AI

Renato Losio: I was actually wondering, as Michael mentioned just before, we still need a human in the loop, but still at the end the human is not taking every single decision. What does good production engineering mean when the humans are not making every operation? Because before AI, I could make mistakes, but I was responsible for every single one of them. Now I'm basically delegating to someone, to AI agents, connector, whatever I'm implementing. I'm just supervising it. I'm in the loop, but I'm not taking every single decision. What is good production engineering at that point?

Michael Hausenblas: I think that in a sense, the agentic flows or agents in general allow us to focus more on what is really important. Think of what is important to the business. Rather than going for low value signals or whatever, I can focus on really what drives the business. I would argue that it very much depends on how an organization is set up. We are operating very similar to what I know from AWS as well. That is essentially you write it, you operate it. The service teams own a service, then you operate it yourself. The SRE team essentially provides support, best practices, audits. The teams themselves own that. I would argue in this sense, it helps the service teams that own a couple of hundred of those services to be more effective, to really focus on what is important to the business and to also have a better way to troubleshoot across all these services if something happens. I think it allows us to focus on what is important. That's the TL;DR here.

Sujana Sooreddy: This is what we were actually thinking about these days as well. When the development has been moved from humans to agents writing code, the first thing that became more important is doubling down on the software engineering principles that we have all worked before. Like verification-first infrastructure, like contract over implementations, checkpoints and automated rollbacks. Always aim for speed of recovery, not speed of prevention because prevention is just not possible in distributed systems that we are building. More and more, when I see that agents are writing the code, it doesn't move our responsibilities of really good software engineering practices, but it actually makes it even more important to double down on it. Like every release has to be a canary promote whole rollback. It cannot be without canaries. Everything has to come. Previously, the metrics, SLOs might have been an afterthought, but not anymore. The SLOs and metrics becomes part of your everyday development.

More and more, I look at it like production engineering in the world of AI is just really good engineering. It isn't much different from what it is. Just make sure that the practices are there and always have the sandbox execution points, make sure the checkpoints are there, work out the different escalation paths, make sure your systems are tiered, which isn't different from what we used to do. That's how I see what a good production system is in this world.

Observability in a Non-Deterministic World

Renato Losio: I think my main problem is that I've always thought until now that SRE was more like, everything was deterministic. I was controlling everything. Now, I'm in a non-deterministic world. What does observability look like when system behavior itself is not deterministic anymore? Do you have any thoughts on that?

Noam Levi: I think it actually also relates to the former question as well. I think the key word is, this is the era of pragmatism. We need to move to a model where we ship code in an evolutionary manner. This is like Sujana suggested about getting signals as fast as we deploy the software, not as an afterthought, but rather as a means to understand how our software behaves in this new revision that we just shipped. The ability to know immediately and to get instant feedback and to wire it to the relevant teams as fast as possible, not just when it's in canary stage, but when it hits production, and get that feedback as fast as possible, as concise as possible to the relevant team. I would argue this is the only way to tackle non-deterministic systems, because non-deterministic systems mean that we need to continuously steer the application through development, and we don't get to stop.

Because there is always going to be an input that caught the system in a state that the system was never in before, and then the system is going to react differently. Also, as I mentioned, the non-determinism is even more severe than I think people actually realize, because what also happened is that people are not just experimenting with their own business logic and ship it much more rapidly. The underlying LLM that now powers this non-determinism is also changing drastically. People are experiencing dramatically different result when they switch from Opus 4.6 to 4.7, from Grok 4.6 to 4.7. Just that notion of what's so-called minor change in the LLM completely changes behavior of what now is a fleet of software that powers finance, powers power plants, powers critical software in many different sectors. The journey to make those systems deterministic is probably going to be a never-ending journey. For most of the teams out there that are not Anthropic or OpenAI, it's more about accepting this truth and create the mechanisms to deal with it continuously.

Michael Hausenblas: Let me start with challenging your initial statement that so far everything used to be deterministic. My analogy would have been anyone who tried to write a shell script that works under any Linux distribution, under macOS, probably knows there is always something that is not there or whatever. I really think that we may now, through the agentic support, actually be able to turn into actual software engineering. So far it was very often compared with gardening. It's more like an art where you do things. Whereas hard engineering, if you think of building a bridge or whatever, it's not that it works all the time, but there are clear expectations. This is a bridge that's built for people or for tanks or whatever. Based on those requirements, you have a certain wiggle room where you can say, the load can be sustained. This is what allows us, I think, to really transition into actual engineering.

If we are very clear with the dependencies, if we're very clear with the requirements. Again, I think that's where the human in the loop is really also necessary, being very clear. Again, you can use things like LLM-as-a-judge to verify things, to use different models, to cross-check. At the end of the day, it requires a human to say like, this is what is the right thing for the business. That is, I strongly believe, at least for the time being, certainly is still the case.

Renato Losio: I got what you meant about the non-deterministic before. That's actually pretty true. I grew up with the entire idea that that works on my laptop, and that's basically the concept. As well, I think that the main disconnect of what you just mentioned, to me as not an expert in the agentic AI space and not an expert on observability, is that, if I look at a bridge, it's pretty clear to me that if you build a bridge in five minutes, it's very unlikely to sustain any load. It's much harder with software at the moment. It's very easy to bring something out and don't have the signal, don't have anything. That difference is less obvious probably. That's, I think, one of the challenges we have as practitioners.

Observability Signals

I was wondering if anything is changing the way we create those signals? Sujana had mentioned before, yes, they have to be designed for a machine, not a human being. Are we going towards more standards? The standards we already had are leading the way or is anything changing? Do you expect any major changes there?

Sujana Sooreddy: One thing which I have seen change drastically is, so if you look at any of your observability systems, like metrics, telemetry related aspects, it's always high input load into the metric system and very little consumption only when there's like an issue versus not. Given that machines are acting on these problems, I think the importance of a really fast way of fetching these metrics, and also, it's an interesting challenge for all of us to think about as well. We actually observed a great level of improvement of results when we are moving from low cardinality of metrics to high cardinality of data. Low cardinality is resulting in more and more hallucinations versus high cardinality on the other end, it's very precise, very deterministic. It is also helping us to actually capture trajectories and then do confidence and probability scores as well. How I see this is, what is changing?

The things that we have made the assumptions that there will be infinite signals and one signal is going to give us an alert where something is going wrong. Now for machines to act, let these signals be as granular as possible and as detailed as possible so the agents could work better. It also comes back to whether it is coming into an era of noise. I think in an AI era, the noise has become more of an issue for us when there are too many signals which have to traverse through multiple different paths, versus very good, like a signal which gives you the complete information about what is happening at that instance, you could drop it fast. For the agents to act fast, just capture that cardinality, that is helping us a lot. That is where I see the change moving forward.

Michael Hausenblas: To come back to your question, I heard two parts, the one around the design for machine consumption and the other around standards. Let me start with standards. There's obviously OpenTelemetry. I'm a big supporter and contributed to that. There are also others. There's the Open Cybersecurity Schema Framework, OCSF, and their interoperability efforts for those. I think that you will increasingly see that observability and anything around security and compliance are merging as we go ahead. The interesting part in terms of machine consumption, and that is something that I'm asking the teams as well. Like, what do you expect in five months or in a year's time? Who will be the main entity that consumes the data? There's a fundamental difference between a human sitting in front of a dashboard, looking at something, trying to make sense of an error rate, and then trying to figure out what service contributed to that, and looking at the logs, for example, to troubleshoot something, and an agent.

That is context. That person that sits in front, at least if the person is onboarded and fully productive, has that context, usually in their head. They know the organization. They can easily find out, does that matter overall to the business, or is that maybe something that doesn't matter too much? Of course, it always matters. If it's the prime revenue generator, I'll probably prioritize it higher. That is something that if we only give access to those agents, what is this operational telemetry data, but not the context about the organization, the product itself, then I think the impact will be relatively not so awesome.

Noam Levi: I think that the word, probably before AI became the most used word of the century, at least in observability realms, I would argue that the word that predates it in terms of usage was context. We were already speaking about context, day after day, morning to night, every second LinkedIn post mentioned the word context is key in some way or another. I think this just goes to show that context was an unsolved issue, even before AI was introduced into play. I think the real issue with context is that it's a double-edged sword. From one end, there is the challenge of knowing what signals are relevant to troubleshoot an issue in production. I would argue that this is potentially like an easy context issue, because for the SRE at edge that needs to recover production, and their reward function is purely production being stable. Potentially, there is a finite amount of context that needs to be accumulated, or as Sujana mentioned, as long as there is recovery baked in into the architecture, at worst case, you can always revert to the last known good state and figure out things without the pressing need.

Context becomes a challenge when we're trying to unleash smarter decisions in other parts of the organization, because then you need to try to ask yourself what signals are relevant in order to answer arbitrary function rewards across the organization. I think this is where AI comes to play, and at least helps tackle this issue, because I would actually argue that the human might not be relevant to accumulate that context, the context will be accumulated and will be determined by the AI. I think humans will operate the decision later. A human will need to have multiple outcomes presented to them, that will steer the current state to a different direction. Two years from now, three years from now, that I think that organizations will still have a problem with AI is the notion of trust. At the end of the day, we're still going to need humans to that very basic need of establishing trust in the decision that has been made.

Renato Losio: You're saying that more than just we want someone accountable, so we want to blame Renato if the production system is down and he was in charge, more as well that we need to be able to overcome that feeling of control that someone is there.

Noam Levi: Exactly. I will give an example that I think most teams already that are shipping code with AI can relate to. I would argue that a lot of teams at least in startups or in companies that are tech first, are shipping most of their code with AI today. It's no longer a question of like, we went through a week without shipping code as humans. These days, probably most of the code being shipped around the world in technological companies is generated by AI. Let's say someone caused a production issue with the code they generated. What is this engineer exactly being blamed for? Are they being blamed for bad code that was shipped to production for not being good enough of an engineer? No, if they are being blamed for anything, if it is indeed their fault, is for them making a decision to ship that code in that state, because no one at this point in time would argue that if they had to write this code manually, the outcome would have been different.

It might as well have been their own humanly written bug, when they ship that code at 1 a.m., because it took them 12 hours instead of one hour of some LLM model with an extra high effort. We're already at the point where accountability mainly revolves around that conversation around trust, and not around the technological prowess of an engineer, or the domain specialty of an employee, if we're going to generalize it for the rest of the parts of the organization, are the focal point of what went wrong.

Safely Delivering Agentic Observability

Renato Losio: You raised an interesting point when you mentioned the importance of context. I want to go back to that one, because it reminds me as well, something that Michael said at the beginning relating to PII and few other things, is, I want to give to an AI agent context to operate safely. I know that I should give the data today. Somehow, I'm not skeptical, but there's that barrier of what I can bring without sharing data that would leak or data that could be processed. The end is, what kind of data you can give to an LLM to operate safely or to make a decision in terms of observability that are meaningful for me, but at the same time, that I keep that data safe enough for the company. Before you mentioned that's a bit of a challenge there in terms of what you can put in that data.

Michael Hausenblas: Again, depending on the environment, like if you think of environments like GovCloud, or you use sovereign cloud, that might be a little bit more challenging. In general, there are pretty good ways to make sure that leakages are not on the table, or if something happens that you have a concrete proof of evidence, and then what happened really after the fact. I'm getting less worried about that part. I can only speak from a B2B perspective, I've never worked in a B2C environment. In the B2B context it might be a specific customer that is really important in a certain region. If I have as an engineer, this context of what that customer means to the business, I make different decisions in terms of scaling something or whatever. It's like, yes, it makes sense, because I want to make sure that this customer continues to have a great experience.

Whereas if I don't have that context, and just look at, this is going through the roof, I need to throttle or whatever, an agent might make the wrong decision. Again, there the trust is really the customer using a certain service, beyond what is in the SLA, saying, I cannot trust your service because of what happened in the past. What kind of context? That is the challenge. I think, overall, it is necessary to provide that context to be successful at scale.

Noam Levi: I think that if I'm to give one practical advice from our experience here, what I think is a big milestone towards feeling confident about delivering agentic features is good sandbox and understanding that security, you need to approach it as a layer rather than something that you can bake by design and as a preemptive measure. I think this is a great opportunity for teams to push for canary deployment practices, having a meaningful staging development environment, ability to mock data or to synthesize close to real world interaction with agentic apps before they are shipped to production. Essentially generating as much entropy and measurements out of this entropy, as soon as you can, as early as you can in the process. Of course, throughout the release to the production itself. We can see it with the frontier AI labs as well, this is also their approach when they are developing their own frontier models.

They are training those models in a sandbox environment. They are working hard on the sandbox, where they get to see in a safe environment, how those LLMs are going to behave. Those LLMs are the ones that are going to power your application very soon after these LLMs leave the sandbox. Hopefully, because we made a decision of allowing them to leave the sandbox, and not as we've seen recently, when they decide to leave the sandbox themselves. When this happens, your agentic applications that are using this LLM needs to be the next one to be put in this entropy cage, where you get to experiment how this app now behaves before you ship it to production.

Sujana Sooreddy: I have a slightly different take on this. For agents to work, the context, it doesn't need to know the internal data, it actually really needs to know the topology and metadata. It doesn't need to know the payloads. It's calling service A, followed by calling service B. Investing in context skills, which are written as decision records, which doesn't have data, but then also invest in a harness, which are like zero data retention boundaries. By default, any system that is written within your organization should come with this default harness added to it, which is zero data retention boundaries, and clear information of what that data is, what should it retain at all in any of its records. It basically becomes as a principle versus an individual responsibility. Individual responsibility is great. To avoid those mistakes, avoid those like leaking huge things, having zero data retention boundaries as part of your system as default, helps in this context. Harnessing this context. Making sure the context is not leaking out, not putting important information in there. For any organization in this era of leaking information, I think that's very good investments to start with.

Model-Level Intelligence Failure vs. Infra Reliability Issues

Renato Losio: How do you distinguish an AI model from making a bad decision from the infrastructure being slow or unreliable? How can a team easily separate model-level intelligence failure from underlying infrastructure or network latency?

Sujana Sooreddy: This is a very big problem that we are also facing. One thing that we are trying to invest in is the trajectory capture. It's like, every agent decides based upon many signals, what is the workflow that it is taking to solve a problem. We are calling that as trajectory capture. Each of these trajectory captures come with confidence scores and probability scores, and they become a base signal before an agent acts. You could think of it like an observability for the agents itself, like how the agents are being run, what is the decisions that the agent is making. So far, we are doing the observability of the systems that we are building. Now it is the observability of the agents, which is like communicating with each other what it is they're working on. Then using eval sets to score these skills that it is calling, the trajectory paths it is taking.

At least this is the bet, of things that we are thinking about, not solved yet. For us, it is like maybe that gives you the signals of what is good, whether it is an agentic hallucination versus an infra thing that is causing an issue. It is more around building the observability and evaluation of the agents and the skills, which can feed into its own learnings than just the observability of the systems that you are building.

Noam Levi: I completely agree with Sujana. I will give our perspective as an observability vendor into this. Initially, groundcover, when we launched the agentic features, it started with an MCP client. Essentially us saying that your local agent, assuming it has access to the data itself, will be able to reason over the data and then will solve the issue. Now this worked to some extent, but the real milestone was when we understood that trying to educate the vanilla agent, let's say your local Claude Code, and to make a decision, a correct decision, without really knowing the platform or how data is organized in the platform itself, gets to a certain limit very quickly. It led us to develop a harness within the platform and move to an MCP interface where, other than have access to data, there is also access to the agent, and essentially giving an agent-to-agent communication between your local agent and our agent as an observability vendor.

The reason this was important, and I hope this also relates to what Sujana mentioned, is that this agent essentially is a harness that steers the agent in the right direction when it makes the decision. This agent layer within the app is a steering layer that leverages the familiarity with the platform in order to provide more accurate decisions at the end. Essentially knowing how to give the scores to which signal is important and which signal is not based on wider understanding of how observability at large looks like. Also, how data is organized and how data is scored internally in the platform depending on how relevant it is to the question the user is asking. Shifting this knowledge to vanilla agents, let's say removing this context of an application, was something that we saw still effective but might hit this ceiling when trying to build true confidence around the decision in the long term.

The Ideal Level of AI Autonomy in Observability

Renato Losio: As AI moves from recommended observability to actually taking action in production, how should we determine the right level of autonomy? We're back to guardrails. Can we really rely on AI to troubleshoot low-risk issues while requiring our approval, so human in the loop at the end, for larger issues? The second part of the question is, AI-based monitoring may cost more as we end up feeding so much data. I think related to automation before. Any better suggestion to integrating AI by avoiding IBL links?

I don't know if the approach of just using AI for low-risk activity and keep the human in the loop for bigger tasks is an approach that you recommend? Do you recommend that direction?

Michael Hausenblas: I would exclude it, per se. As I said, the actual decisions, if it's something that might impact the entire business, is something where at the end of the day, you want to have a human accountable name attached with that. It also ties my head a little bit in with the previous question, and I think Sujana you mentioned it in the beginning in terms of you can't really prevent things, you can only embrace the failure. We typically do things like chaos engineering that might surface also certain things. Does it come from the model? Is it some underlying infrastructure issue or whatever? If you have these practices in a continuous manner, including game days, fire drills, the trust builds. The people operating and then depending more and more on these agents can put more trust into them. Also sponsors or business representative, stakeholders might be more comfortable with that. I know this touches pretty much everything. The TL;DR is really, you cannot prevent things, you can surface them and then assess the risk based on those findings.

Noam Levi: I think that deciding if you want to give autonomicity to your agent, this confidence, it's not a question that revolves around the agent, it revolves around the organization. You can already probably in many sectors, give AI the ability to be autonomous in some parts of the organization. It really comes down to how resilient your organization is to having flux or to having deviations in that specific area. How important is it to you, to your organization to take risks in order to move fast, because maybe your organization, your main competitor decided to go on a no holds barred when it comes to velocity and just ship as many features as possible, which also put you with your back against the wall in order to ship at the same velocity. You might not be able to allow yourself to play it as safe as you would like to.

There are many applications there that are not power plants, that are not nuclear reactors. Let's talk about sales organizations, which I think is an interesting example. A lot of potential prospects this day for many technical sales representatives have rules in their agents that will filter out outbound calls, unless they have specific keywords that are of interest to them. At this point, or when you don't know exactly what kinds of rules are in play, and you can see it already with many solutions in this area, salespeople are now trying, and they are very smart at doing that, trying to experiment with autonomous agents that will do outbound prospecting and will figure out success rate. I think that's a great example of where it is almost potentially inevitable to adopt autonomous practice to AI, because this ability to experiment fast is instrumental for your success in your specific domain.

I think this is relevant for many other domains. Maybe another question to ask yourself is not whether I am willing to give these specific agents the ability to be autonomous. It is to ask yourself, have I as an engineer or as an organization created a sandbox for this agent to make mistakes, where I feel comfortable with having mistaken this specific domain in this specific playground that I created for the agent.

Getting Started with AI (Action Items)

Renato Losio: I'm not saying that 100% of the code we ship now is AI generated, but definitely we are more comfortable at the moment in shipping code that is AI generated than having an AI agent managing our production, observability, and SRE. I see much more adoption in code development. If I am an SRE or if I'm a team managing the observability platform of our production environment or whatever, that has no AI in place today, what should be a simple first step in that direction? You already mentioned just start with a sandbox or whatever. It's like, what can I do as a very first step into that direction to try it out without feeling overwhelmed by everything changing? Thinking a bit of a task as an action item, after the roundtable, what can be something more that I can take action on?

Michael Hausenblas: Maybe a good place is to start in a greenfield environment rather than brownfield. Maybe something small that you start from scratch. Brownfield environments where a lot of existing dependencies and whatnot exist might be a little bit too much as a first start. I think it's really a matter of getting into that and being very aware, being clear with the expectations. It's not a silver bullet that will change anything and everything, but you need to find these feedback loops to build trust both with the agent and in the organization and ensure that you can actually not just trust, but literally verify all the outcomes, which if you look at UI or whatever, it's not an easy thing to do.

Sujana Sooreddy: For an SRE who is working on systems, I would start as writing down the actual jobs to be done. When X happens, I want Y so I can see Z. Whatever they're doing on an everyday basis, for any of the alerts, for any of the runbooks that they have, just write it as like, what are the actual jobs that they're doing or jobs to be done. That itself is the base context. For AI, the most important is context and the context which is written in the most deterministic way. I find like for SRE, the best way to start today is like just write what are your jobs to be done in a way that the machine can understand. Then, on top of that, you could just then attach, like the Google SRE has these levels of AI autonomy framework. It has like L0, which is manual execution, goes to L4, full autonomy. For each of these jobs, mention how much of an autonomy it should be. That is where, if I'm starting tomorrow, I would just put that context in.

Noam Levi: I think that at least I can tell people here what worked for me recently, when I also started to explore agentic adoption into other realms of what I do in groundcover. I try to connect it to every single digital footprint that I have, whether it's Notion that we use for like a company brain. I let it search my Slack messages. I connect it to my email. Then I simply ask it what in my day-to-day job in the last two weeks can be automated, can be made easier. I think that the best way to build confidence around anything is to start with low-hanging fruits that are always there. Nobody is only left with only hard challenges. I would say for many people, we are all lazy about some aspects of our work. You start to build confidence around automating these lazy parts, this will give you confidence with working in that specific way, as a first step. From there, take the leap of faith of automating the more complex stuff. It all starts with just building this digital footprint of what are you doing, and let the agent familiarize with the day-to-day tasks that you are making.

 

See more presentations with transcripts

 

Recorded at:

InfoQ Live is a virtual event designed for you, the modern software practitioner. Learn from extraordinary speakers driving change and innovation.

Oct 01, 2026

BT