Why the Gateway Approach to Agents Sucks (and What to Do Instead) with Marek Poliks

Share via Twitter Share via Facebook Share via Linkedin Share via Reddit

Get more video from Redmonk, Subscribe!

James Governor sits down with Marek Poliks, Head of AI at LaunchDarkly, for a spicy conversation about what it actually takes to manage agents in production. Marek makes the case that feature flags were almost designed for this moment: pulling agents out of application code lets teams gate, roll back, and iterate on whole agent harnesses as testable variations rather than hard-coded entities. Along the way he argues that everyone is now experimenting in production whether they admit it or not, and that the popular gateway approach introduces brittle third-party dependencies at the most security-sensitive point in the stack. According to Marek, observability “will not save us.” Passive logging is a lagging indicator; the real value lies in runtime control and adaptive triggers that intervene before bad things happen. The discussion ranges across progressive delivery, software factories, healthcare and finance adoption, and closes with Stafford Beer and outcomes-driven engineering.

The RedMonk video is sponsored by LaunchDarkly.

Links

Transcript

James Governor (00:04)
Hey, it’s James Governor. We’re here for another MonkCast. I’m really pleased today to have Marek Poliks head of AI at LaunchDarkly. And we got some content that’s maybe well can’t actually why did I call it content? It’s not content. We’ve got content. Content is close to slop. What it is, is a conversation with value. Yes, not content. It’s not content, it’s a conversation.

Marek Poliks (00:17)
There’s your first mistake.

There you go.

James Governor (00:31)
and it’s it’s possibly a little bit spicy today. so I think that if we if we had a title, and and you know sometimes the title comes later, but it’s it’s something something like having having done a little bit of planning with Marek, it’s how the gateway approach to agent management sucks.

And why observability and why observability is not going to save you. So there you go. Now, a lot of you know LaunchDarkly from the feature management space. Maybe you don’t even know that they’re in this agent management space. certainly what I’ve seen, lots of vendors, as I say, coming in from kind of

Marek Poliks (00:51)
Yeah.

James Governor (01:11)
You’ve got API gateways, now being AI gateways, now being agent gateways, you’ve got messaging vendors in there. All the security vendors are all like, yes, we are gonna secure the agents, we can help you manage that. so I was kind of interested when I was talking to Marek about this. We’ll we’ll get to that during the conversation. But I guess from my perspective, as I say, we very much know LaunchDarkly as a

feature management company. You invented the category and and then I I guess what’s a nice feature management company like you doing messing around in the AI agent space Marek? And welcome to the show by the way. Welcome to the show.

Marek Poliks (01:41)
Yeah.

Well, I’ll say… Thank you. Thank you so much for having me. I wore my red shirt today so that I could be extra spicy, I guess, in terms of the opinion setting. Though it’s matching the color of my nose from a lot of nose blowing. have a little bit of a cold, so apologies if I sound bad. Yeah, so what business does LaunchDarkly have in the AI space? I’ll say I joined LaunchDarkly because I came to the same conclusion that a lot of folks inside of LaunchDarkly and…

product and also a lot of launch work these customers were coming to, which the idea that like feature flags are actually really, really useful for agents. So as someone who was building a lot of agents, you know, a year, year and a half ago, certainly earlier than that as well, I found it to be really useful to take the agent outside of the application space itself. You the application space, you have this world that you’ve defined where you’re really looking to do specific things. The agent that lives within that world or in many cases the

16 agents that are living with that application space or 80 agents that are living with that application space like if I think of that Application space as one topology as like one neutral world. It’s really difficult to govern those agents It’s really difficult to shut those agents off if they’re doing something problematic It’s really difficult to root cause those agents if they’re doing something problematic and I think most critically it’s really difficult to like improve on those agents because those agents are kind of like

hard coded entities that are living in code. So prior to me joining LaunchDarkly, I was starting to put agents behind feature flags so that I could iterate on them. I could like swap in an agent, swap out an agent. That was definitely the most important thing early on. But then I also learned that like, well, the way of improving an agent, if I start with an agent, let’s say I built an agent harness, I’m like, all right, this agent harness is good, but I’m gonna make it better. If I improve that agent harness by just,

iterating on like a single surface, I’m doing something really weird in terms of like a non-deterministic system, because every time I’m testing that non-deterministic system, I’m basically rolling the dice again. But if I’m just testing one version or variation of an agent harness and then seeing if it gets better, seeing if it gets worse, I’m kind of playing an awful stochastic game with myself. And the better thing to do is actually work with a large kind of communion of variations of agent harnesses, test them out at some level of scale.

James Governor (03:46)
Mm-hmm.

Marek Poliks (04:09)
and then look at the probability distribution of those performances. And that is also something that’s really useful. Well, I can keep going on why I feature flags are great for agents, but like taking them out of application space and putting them somewhere I can see and think about them, and then like treating them as like these variational kinds of things that I can test at some level of scale were two really kind of like no brainer conclusions for me that led me to join LaunchDarkly and start building out what we call agent control, which is principally this.

super complex and think really robust way of ultimately delivering agents behind feature flags.

James Governor (04:43)
Okay. So I I think one of the points you’re making there is that we sort of live in an age now of of of experimentation. I mean w in this agent world, kind of everything is experimentation, something we’ve spoken about for the longest time. You know, we would have feature flags, but how many feature flags? And I guess now there are so many flowers blooming, that, you know, we need to be able to I guess nip those nip some of the nip nip nip gardening.

Marek Poliks (04:52)
Absolutely.

It’s gardening, right? Yeah.

James Governor (05:10)
ni nip some of them in the bud and as as you say, it’s it’s much more probabilistic. So is it is it partly that that just the the degree of experimentation aligns itself with what we’ve been doing with feature flags already?

Marek Poliks (05:26)
100% I think that’s something that we keep kind of joking about internally is like, maybe like this is actually the best usage of feature. Maybe like this is what feature flags were designed to do all along. And we’ve just kind of been anticipating this moment because like we were, moving to a place where, you know, first there was a lot of experimentation and kind of pre-production context. Then the feature flag came around and allowed you to do some experimentation and production context. Then agentic AI and you know, AI in general kind of had its big surge.

And now everyone is experimenting in production, no matter what they say, like absolutely no matter what they say, they’re doing it all the time, all the time in production, because no matter what you’re rolling dice and customers are involved in those dice rolls. And so the question is like, whether or not you’re collecting information about those dice rolls, or just casting dice all the time. And so yeah, for us, like it’s very much a moment of taking advantage of the fact that no matter what you’re experimenting now, everything is experimentation now.

James Governor (06:01)
Mm-hmm.

Marek Poliks (06:21)
It’s just a matter of actually ensuring that you’re building a proper experiment that’ll teach you something as opposed to just gambling I guess.

James Governor (06:28)
Right. And so on that note, I mean, i is this so we’re all we are all now fully testing in production. That’s that that ship has sailed. Is that is that the articulation?

Marek Poliks (06:39)
Yeah, I mean, I don’t think that’s a controversial perspective at all. And like to a certain extent, we always were, right? Like there’s never been a place where we like to think of compute as discrete state machines. We like to think of like moving through a series of gates and like everything being deterministic and determined from the beginning to the end. But as soon as you put anything non-deterministic into a loop, including a human, you know, the OG non-deterministic system, you get an aggregate non-deterministic system.

James Governor (07:05)
Yeah.

Marek Poliks (07:08)
And that’s ultimately what happens. As soon as you allow for the concurrency of deterministic systems with our beautiful non-deterministic world, you get a non-deterministic space. And that necessarily incurs something like experimentation, right? In that you’re testing alternatives, you’re testing possibilities, you’re opening up vulnerabilities in a lot of cases. That’s always been the case.

We’ve just kind of ratcheted that up like mega style with agents because now it’s both sides of the loop that are highly non-deterministic. No matter how much we want them not to be.

James Governor (07:45)
Yeah, I mean distributed systems, by their nature, we may w it want and hope that they are deterministic, but we certainly can’t rely on that. things break. you know, I I think the I always think of of of Mike Tyson, you know, everybody has a plan. Everybody has a plan until they get punched in the face. so, you know, people had pre prod, but prod is liable to punch you in the face.

Marek Poliks (08:08)
Yeah.

James Governor (08:15)
I think everything everything can move that way. So yeah, I I I don’t think it’s a controversial statement in in 2026. So one thing I think that’s interesting is when LaunchDarkly started to talk about runtime control. I had issues with this. I’m like, what do you mean? LaunchDarkly does not do runtime control. For me, runtime is it’s about going down a level.

Marek Poliks (08:31)
Mmm.

James Governor (08:42)
You’re working more at the the feature level. So how is that runtime control? And slightly annoyingly, as we see so often in this era of AI, we have to reassess our priors. And and during this conversation, I I think that that kind of seemed to me that that yeah, like at least in this new world,

The way that you’re using flags, I think I’m gonna say it’s reasonable. I look how unhappy I look saying it. I think it’s reasonable to say that you’re doing runtime control. So talk a bit about that transition for LaunchDarkly. Because again, I don’t think that’s something that people think of you for.

Marek Poliks (09:11)
Yeah. You’re 100 % right. I think a lot of people, including some customers as well, think of this as a kind of binary gate that’s outsourced from the application. You have something going on in the application space. You’ve identified that this feature, this component, you’re to take that out of the app. You’re going to move that to the LaunchDarkly ecosystem. And then you’re probably just going to turn it on or off based on certain conditions.

most classic possible example is like, we’re going to put out a new feature, but we’re just going to put out that new feature to beta users, right? And if it’s bad, we’re going to shut it off immediately. And that’s like a very kind of classic LaunchDarkly use case in that you’ve taken something out of code, you’ve moved it to another place where you can conditionally gate whether or not it’s active based on customer segmentation. So whether or not a customer is part of a beta user segment, and then you can shut it off really easily by toggling it on and off from the outside of the application logic.

Which is great. So what that does is that decouples deployment from release. Now I’m just like straight up spitting LaunchDarkly marketing speech, but I think it’s actually pretty good. Yeah. Yeah. Hell yeah.

James Governor (10:26)
I I I honestly I I I grew up on this stuff. I mean I’m I I I wrote a book about this stuff. I mean so

we’ll get to progressive delivery in a minute, ’cause you know, I mean I don’t mind you like being spicy about everything else, but if you start coming after my baby then this is gonna be so anyway, keep going, keep going. But yeah, no, that world definitely makes sense. That is the expectation.

Marek Poliks (10:49)
So where does runtime control come in? So let me just articulate the decoupling of deployment from release deployment, of course, to the uninitiated means the moment that software is ultimately actualized, you’ve got your servers, you’ve got your clusters, everything is kind of rolling out into the world and release is being separated from that moment because you’re actually turning on a feature or turning off a feature after the fact, right? After the software is already out there, you have the ability to release some new piece of software or unreleased some new piece of software.

James Governor (10:55)
Yes.

Marek Poliks (11:19)
based on live conditions. Now live starts to in here just a little bit. This gets really, really exciting with agent control in a lot of ways, but also just kind of generally how LaunchDarkly has progressed over the past couple of years. But let’s stick to agents, because I’m an agent guy. think of it this way. Let’s say you have a new version of an agent that you want to roll out. Let’s say you’ve got agent A, you want to move to agent B. You’ve done some pre-production, offline evals. You’ve validated that agent B is good. So we’re going to roll it out.

The first kind of like primary primitive that LaunchDarkly lets you use is what we call a guarded progressive rollout. And what that ultimately means is I can gradually stage that release from A to B of an agent harness, again, separated from code, just the harness, the prompts, hyperparameters, policies, tools, model, model provider stuff. So just that is on its own release track independently from anything in the kind of application body. And you then have some like live signals.

but we’re SDK driven. Now you can hear, I’m gonna start to ramp up like the dissing. I’m trying to think of like a safe way to say it on air, but like the negative opinion I have towards like strictly gateway based logics here. But since we’re, since we’re, okay. Just talk to the air the hell yeah. Let’s go for it. Good, good. So,

James Governor (12:33)
Wait, I I think I think it’s I think we’re all friends here. I think if you wanna say if you wanna say shit talking, I think we could just about I I think we could just about get away with that, here.

Marek Poliks (12:48)
So we’re inside an SDK. So we’re inside the actual application space itself. So let’s say application initializes. It pulls a series of flags that are available via LaunchDarkly. You serve 50 trillion or 60 trillion, I forget what number it is, but it’s trillion. And I love being able to say trillion every day. Like lots and lots of flags every single day. Lots of points of presence all over the globe. Those flags are retrieved. Yeah, exactly trillions.

James Governor (13:09)
It’s the doctor the doctor the Doctor Evil thing, right?

Marek Poliks (13:17)
So application initializes, grabs the flags, and then inside the application space, this is where we get towards runtime control. We’re going to switch that trigger from A to B over some progressive scale. So I’ve identified, let’s say, my enterprise customers. I’m going to roll out from A to B of this version of an agent. And I stage that progressively. So gradually, it’ll get switched on to those customers. And then here’s the important bit. There’s all kinds of metrics that are living in that application space that can be attended to.

The application space has variables and all those variables can be emitted to LaunchDarkly or leverage inside the actual application space itself as some kind of tool for what we call runtime control. So let’s say you have an online evaluator that’s sampling some percentage of exchanges for accuracy or bias or contextual grounding. Or let’s say you’re polling the model response API and you’re getting a sense of how latent it is or whether or not you get any errors, et cetera.

James Governor (14:07)
Mm-hmm. Mm-hmm.

Marek Poliks (14:15)
That information can be leveraged two ways. Let’s say you got a bad series of bad online eval scores. You can automatically roll back that progressive release from A to B, blam, done, rolled back. You don’t need to page an engineer. You can page an engineer and let them know the rollback already happened. They’re good to go. But let’s step further into runtime level control. This is where I get really excited about adaptive triggers as we call them. But let’s say you have some level of latency that’s been determined as unacceptable.

James Governor (14:31)
Mm-hmm.

Marek Poliks (14:42)
Or let’s say you have some level, let’s say an online evaluator has deemed the response of a model to be unacceptable in the context of the actual application space itself. You can immediately escalate to another variation of that flag. Because we’re serving different variations of flags, because we expect you to go through the vetting process of flags that are going to be exposed to the production surface. Let’s say you got a bad review from an online evaluator, you can escalate to a more powerful model. Or let’s say, you know,

James Governor (15:04)
Yeah.

Marek Poliks (15:11)
Anthropic is timing out again for some reason. You can escalate to a more powerful either connection or to a different model provider or to different model altogether. And that’s inside the actual runtime logic. That’s literally at machine time.

James Governor (15:30)
Okay. So th therefore runtime. So here’s the thing. And what what I do find interesting about this, I mentioned progressive delivery a minute ago. yep, me and and and some friends wrote a book about that. feat flags was definitely part of the underpinnings of that. you know, we we needed observability for the feedback loop, really leaning into obviously automation has come a long way.

Marek Poliks (15:33)
Therefore runs out. Yeah.

James Governor (16:00)
You know, if think about CI C D, a lot of that stuff was you know, really pre-cloud, so cloud abundance wasn’t part of that world. we’re certainly, whilst we’re worried about token economics right now, there is a huge amount of abundance of infrastructure that we can take advantage of. How do you get alignment between the business? How do you get live in a world where developers have autonomy?

Marek Poliks (16:12)
Okay

James Governor (16:26)
and where engineers have autonomy so that we do effective work. And we just felt there wasn’t an articulation for all of this stuff that that captured all of those things. And so we talked about progressive delivery. Now, I think to your point, we started writing the book pre-Chat GPT. And let me tell you, when the world is changing as fast as it is, you’re kind of like

Does this book have any relevance in this new era? So we were a little bit worried about that. But but I I think what we’ve seen is that that clearly it does. And I I think that in of of of what you’re talking about, yeah, AI and using models and agent based development very much need some of these concepts around as I say

Marek Poliks (16:56)
Yeah.

Progressive delivery has never been more important. I mean, like it literally, I can’t imagine it being more important than it is at this particular moment.

James Governor (17:25)
Yeah, and I think what’s interesting there is having leaned so heavily into autonomy, in the core concept, we weren’t even thinking about autonomous agents. We were thinking and and you know, and then alignment, we were thinking about alignment, human alignment. And obviously now, once we start getting into loop-based development, alignment

Marek Poliks (17:35)
Mmm.

James Governor (17:48)
That is really critical. So you’re sort of seeing these things, you’re like, okay, we’re just describing this new world. and I think what’s really interesting is of course that transition you’re making. now, loop-based development. Like that’s you is that is that something is that something you’re seeing i in your customers? another word I think people are beginning to talk about, like you know, software factories.

Marek Poliks (18:08)
You

James Governor (18:16)
Like, is that well, I mean you’ll launch Darkly, so they’re definitely dark factories. But well w one of my questions is I’m certainly seeing and and look, I mean it depends from org to org. You know, I could go in, I I was lucky enough, I was hanging out with Richard Sworder, who runs product at the Atlassian Williams racing team the other day, and he’s definitely all about look,

Marek Poliks (18:17)
Dark factories. Yeah.

James Governor (18:44)
we’re gonna we’re gonna build stuff, we’re kind of not gonna mind if it breaks. What we are gonna have is really good pipelines and the ability to roll things back. But you can’t hear actually. I mean, I used that the the we you know a few minutes ago. We were talking about the sort of like flowers blooming and so on. They talk about leaves rather than branches, right? Things that you could snip off, that is fine. now they’re pretty advanced, obviously, sophisticated and they use the technology. My question for you is a

It are normal humans, your your sort of your normy customers doing this stuff? Or is it just sort of customers south of market that are like, we’re living in the future, you know, token maxxing, completely AI pilled? What are you seeing in in the sort of those the I guess those different modalities, different customers? Are real customers living in the future with you?

Marek Poliks (19:36)
Yeah, absolutely. think as with all things, like it’s less of a, let’s say a movement between, I don’t know, two separate sides or something like that. And then an absolute tear in the fabric of like how people do software. You’re either all in or you’re unbelievably behind, but the actual constituencies of those sides is

crazy. Like it really doesn’t make any sense to me. It’s not split based on just like AI native versus non AI native. It’s not based on industry or even like degree of, you know, subjectedness to regulation. It does seem to be pretty random and largely based on leadership principles and a little bit geography. I guess I’ll be real with that geography is like pretty, pretty leading indicator there. On one side.

James Governor (20:22)
Are you are you saying that Europeans are slow at AI adoption?

Marek Poliks (20:27)
Well, you know, there’s this act that keeps getting pushed back every six months that seems pretty threatening. And I’ll tell you that, you know, my customers in Europe are justifiably very concerned about it. And at the same time, I guess just shamelessly, say I really I get really excited reading the the EU AI Act because of the just degree of conformity to what it is that we offer at LaunchDarkly. Credibility, traces, humans in the loop, you know, an important decision making, apertures like.

James Governor (20:30)
Right.

Yeah.

Marek Poliks (20:57)
Anyways, but yeah, I there’s certainly a little bit of that for sure. But I think the thing that I’m most interested in is in how healthcare and financial services are actually leading the way in a big, big way. This is where I see actual full scale software factories starting to be built. And moreover, this is also where I see, I guess, the most surprising moment of my career at LaunchDarkly thus far was with a highly regulated customer who was working with, let’s say, an OpenClaw like system.

already starting to build fully autonomous architectures here. I was really surprised to hear that because there’s a lot of security vulnerabilities and problems when it comes to managing fully autonomous agents. But I think there’s just a sense that this is happening and that maybe the posture or the approach to things like security, data provenance, observability in general has a lot more to do with exactly what you’re describing, this kind of leaf-based approach where

you’re kind of continuously pushing out new things at the end. Those things can be rolled back easily and safely, and then kind of you can continue to progress forward. Like lots and lots of like kind of filigree that’s growing on the outside of these teams. But at the same time, like the more substantial stuff, I haven’t quite seen that yet where you’re seeing like, you know, full like large scale institutional architectures being completely automated that I have not seen outside of, you know,

the folks at the bleeding edge of the industry, but everyone gets the sense that it’s inevitable. The number of conversations I have with people who say, you know, the workflow is dead is a lot. I mean, that’s not an insubstantial amount. And then on the other side, you have people who who have spent their past year and a half, getting their data story in order so that they are ultimately ready for the agentic turn. And I think they’re learning very quickly that that last year and a half was full of potential lessons that they haven’t learned yet.

James Governor (22:37)
Yeah.

Marek Poliks (22:56)
So that’s a zone of real vulnerability, I think, for a lot of folks.

James Governor (23:02)
Yeah, I think I think there’s a lot of truth to that, but on the other hand, there there is a lot of prep to be done. I mean, you know, it’s one thing like get the data foundation right, but from a context perspective, really ensuring that you you want your agents to be working with the right context. And so, yeah, talking to an organization recently and and quite literally, you know, they they turned the agents on, the agents had gone out, started getting access to things, getting pulling data in.

And they were getting like fifteen years of SharePoint data. And it it turns out that the model was not. Yeah, no, that was a serious yikes moment. So you know, definitely take a view of of cutoff points, you know, like that stuff may not be relevant. I I remember I I thought it was funny when you know Amazon first started talking about training models and and you know how it was gonna improve life of the Amazon customer. And they’re like, well, you know, yeah, we have

Marek Poliks (23:35)
Yikes.

James Governor (24:02)
you know, all these all these years of of like Amazon documentation, we can train the model on and everything. I’m like, you don’t want to feed the model on old docs. That’s like not something that’s gonna give you a more effective context. I mean, unless you’re working in a very specific sort of modernization con So I I do I I I I I think where you’re absolutely right and what we are seeing is

Yeah, waiting around is not a good idea. that

Marek Poliks (24:32)
Yeah, you have to kind

of do both. I mean, totally context is critical and important. But if you’re like, you have to learn the lesson, like that lesson that you described of that, that person that you spoke to, that’s a lesson you learn by actually leveraging an agentic system and seeing what kind of context is actually pulling. It’s critical to do both at the same time.

James Governor (24:51)
Yeah, a hundred percent. So so we’ve got, as you say, sort of healthcare. and there’s you know, there’s the on the healthcare thing and regulated in industries and yeah, I mean now I’m not saying that the entirety of every organization is working in this way, but I’m not talking to many enterprises that don’t have pockets of teams working i in this way. And to your point about like product velocity, it’s like well

Marek Poliks (25:14)
Exactly. Yeah.

James Governor (25:20)
Obviously disruption continues to be a thing. but but you know, I I I don’t know, I think it you know, Charity Major says, you know, five nines don’t matter if the user isn’t happy. Well, you know, you know, being fully industry compliant really doesn’t matter if you go out of business. I mean, the the the I’m a big believer in in the value of of regulations, consumer protections and so on.

Marek Poliks (25:48)
course.

James Governor (25:48)
But pretty clearly the the the the challenge for organizations today is they’ve got they’ve they as you say, that you are at such a disadvantage if you are not moving fairly quickly to a new way of of of of building software. So I think one thing that that is interesting just about this whole

This willingness to say to the machine, you do the work, we you know, commander’s intent, right? This is what we want. Now you go and and and build the thing. is it in a way and like so software factories, we’re talking about that now. That’s not new, is it?

Marek Poliks (26:34)
at all. mean, it’s in it. AI is just an intensification of automation, right? Like, I mean, it’s really just it’s like a really nice kind of linear growth of automation applied to pectoral muscles to machine controls to fully automated systems now to the automation of automation itself, right? Like, it’s really kind of nice curve there. Totally, I completely agree. And also in the software space, it’s also like absolutely not a new thing.

James Governor (26:36)
I mean that’s that

Marek Poliks (27:03)
If you think about what even something like CI is like, I mean, that is like just the progenitor to this and in absolutely every respect, it’s just the introduction of automation, the introduction of loops. I think kind of most importantly, an increased attentiveness to outcomes as opposed to inputs. Like right now, I think like we’re starting to see this across the industry in general is like actually really alluding to what you were just saying around something like regulation. Like there is a heightened attentiveness to outcomes.

James Governor (27:19)
Mm-hmm.

Marek Poliks (27:31)
the actual customer outcome, the impact to the P and L, like the impact of the business. Like these are outcomes that are measurable and that are ultimately the kind of like, you know, the thing that one steers toward. And so there’s less of a focused on less of a focus on inputs, more of a focus on outputs, but that’s been a progressive component of software development for a very, very long time. Definitely.

James Governor (27:51)
Yeah,

I’m a absolutely right. I mean, you know, and I I do think that that’s it’s I mean, at any given time and we you know, we were talking about the transition to product as as the key way of thinking. Sort of moving away from this focus on projects, which is what we used to do. and again, a lot of that was, you know, input based, to to very much ’cause products have by their nature are about outcomes. It’s about value

Marek Poliks (28:03)
Exactly. Yeah, yeah.

James Governor (28:20)
to the user. And on that note, let me let so we I wanna I wanna rewind a little bit, because we we we said we were gonna get spicy. I wanna make sure I wanna make sure that we get spicy. So you explained you you you you explained why you you you you’ve got some views

Marek Poliks (28:33)
Yeah.

James Governor (28:45)
on why the gateway approach doesn’t fully make sense. Is some of that to do with th the I guess the number of black boxes that we currently face? Like what’s the what what’s wrong with you know, securing endpoints? What’s wrong with the endpoint approach?

Marek Poliks (29:07)
Also, there’s nothing in principle wrong with a gateway. And in fact, I think every mature enterprise AI body should have a gateway. That is a critical kind of control point. Yeah, some of my best friends are gateways. Absolutely, absolutely. But I mean, they also introduce a lot of issues. Especially, so first, especially if you’re using a third party gateway, this means that you’ve introduced a serious level of vulnerability, a serious level of dependency.

James Governor (29:16)
You are you saying that some of some of your best friends some of your best friends are gateway? Yeah, yeah.

Marek Poliks (29:36)
critical juncture point within your system. And this tends to be how like a lot of AI observability and a lot of like AI tooling, especially around governance is ultimately instrumented, including things like guardrails is like by introducing a third party dependency that adds latency, that adds, you know, single point of failure logic right at the API call itself to the model provider, which is already such a sensitive, like, you know, infrastructurally contingent like moment that’s happening applications right now. And I think like the

The bigger question there is, so like if you’re sending information, that’s critical. That might go to, let’s say, the enforcement of whether or not someone has access to a model or the enforcement of whether or not a guardrail should be imposed. If you’re sending that to a third party, that means you’re sending your customer contacts, so everything the customer sends in in the form of like a user prompt, you’re sending the model’s response, you’re sending like all of this,

James Governor (30:22)
Right.

Marek Poliks (30:36)
business critical PII forward security rich information through a brittle third point of failure. And to me, that’s like a serious problem. And so the majority of people that I see, especially like the advanced folks who are working in highly regulated, when they’re building gateways, they’re like confronting this impossible problem, which is like, how do I regulate what is going into and out of these models without looking into what’s actually being said without storing any of that information anywhere because I’m not allowed to.

And that becomes a really, really, really complicated question. And so you can certainly build some instrumentation around that, but that instrumentation starts to look a lot less like a gateway and a lot more like runtime control from our perspective. But yeah, so just to kind of put a button on it, like gateways, like certainly centralized administration of AI is a good thing. If that centralized administration doesn’t have an understanding of the constituent components of the harness of a given agent, it can be kind of toothless, like…

Most gateways are just like, have access to this model, you don’t have access to this model, or like, maybe if the model starts to suck, we’ll switch to this model. But like, they’re not a like highly active control point, because the amount of context that’s actually being handled there is actually not very rich. You don’t have the full harness information often. You don’t have necessarily like a tools registry or a skills registry or anything like that in place that you can actually supervise. You’re just working with an application that is a client that is like somewhat invisible to you. You have an API call that you’re handling.

And that’s kind of it. And so like you’re limited in terms of what you can control, you’re limited in terms of your governance, and you’re also sitting at like the most contingent, the most brittle, the most like security complex point of the entire architecture. And so that for us is like where it’s kind of cooler to be inside the application, where we can do things like provide guardrails and even online evals and other kinds of metrics from inside the application without necessarily revealing any context back to LaunchDarkly at all.

because we’re serving harnesses that go to a place that can be enforced in application logic to do certain things. Like for example, our online evals work by sending you a harness and saying, do an online eval. They don’t return any information to LaunchDarkly. There’s no API call to LaunchDarkly that’s being made in the middle of the run. No, it’s all happening nicely in machine time. No added latency, no requirement to pass back customer context, for example, or customer query. LaunchDarkly, and at the same time, you still get your eval, which is nice.

James Governor (33:02)
Okay, okay. More spiciness. So you’ve just talked quite a lot about sort of instrumentation, the value of climate tree but I mean you do it will observability will it or will it not save us? So you’ve you

Marek Poliks (33:18)
It will not save us. Observability won’t save us. It’s a lagging, observe. Can you think of a worse word? Who wants to observe a dynamic, incredibly contingent, powerful system? Observability to me means passivity. means like looking at a giant log of like every bad experience my customer’s ever had. And those experiences have happened, right? Like that’s what it means. It’s like living testimony that something bad occurred. And the goal is to get ahead of that, right? And that’s even more important in the agentic era because like,

It’s the agentic era, like real bad things can happen. The more useful a system is, the more critical, contingent, complicated information it has access to, right? The more agency it has to do things that are potentially bad. The blast radius is large already. You see like that blast radius can have a serious impact on brand integrity, for example. It can also have, you know, lots of complexities when it comes

James Governor (33:51)
Mm-hmm.

Marek Poliks (34:13)
things like security, for example, of course. So the problem with observability is that it is lacking. And so that’s why we’re so interested in runtime control. We’re interested in doing things inside the turn of the conversation. We’re interested in giving customers the ability to intervene before the bad thing happens, right? That for us is really critical. That doesn’t mean that information is bad. Information is great. It’s super important to have information. That being said, I can imagine, absolutely imagine,

a year from now where we do see software factories really humming along, where that information becomes significantly less useful. Already, I mean, I’m sure all of us have been in situations where we see observability information about a system that doesn’t exist anymore, right? And like, that’s great. All right. Like, I get some information about something that no longer is part of a chain that I’m actually leveraging. And as we see, like rates of rollout to production move to the…

every three seconds, great, like, I mean, just continuous, continuous, continuous software changes, that raises a huge hurdle, not just for things like observability, but for things like versioning. Like, it just like creates this like huge question of like, what is this software thing? And what about its behavior is actually accessible to me as an observer, as someone who wants to learn about it, as someone who wants to intervene in it. But yeah, I think.

I observability is in a really kind of challenging position. That doesn’t mean that like logs are bad. Of course, every logs are great. But what it means is that you need more. Like you need the ability to actually intervene. You need the ability to get actually active inside of runtime. You need the ability to afford problems from actually happening. And to me, that is more useful than that information about a thing that’s happened that may or may not be reproducible ever again.

James Governor (36:05)
Yeah, you’re you’re you you’re definitely you’re just you just do not Yeah, I mean that that’s that’s quite a spicey thank about observability. I’m not wrong. observability is all driving along, looking in the rear view mirror. I mean, you know, I I’m I I I I think that that the Yeah, that’s quite a log based view of observability, that’s what I would say.

Marek Poliks (36:12)
No, no, it’s absolutely not

I think that’s probably true. Yeah, but I mean, I think like, to me, the more useful and interesting component of observability, let’s take it up like the higher level perspective. And the last time that we met, we chatted, we nerded out briefly about like the history of software factories or dark factories in general. And I was thinking about Stafford Beer, so early cyberneticist, really important for the phrase, “the purpose of a system is what it does.” Posse with great, great simple statement.

Understanding what a system does is really important. Information around what a system is actually inducing in terms of customer behavior is the most important thing as we shift to fully outcomes driven engineering, like a thousand percent. And it’s that particular regime of observability that I’m unbelievably committed to and interested in. It’s significantly less, however, about the understanding the, let’s say the superposition of states.

that an application is in at a given point in time. That to me becomes a lagging indicator. Instead, like the focus just continues to drift to the right, continues to drift towards the customer, and continues to drift towards like actually motivating the desired behavior of the factory without too much of an ability to actually get a sense of what is happening on a millisecond to millisecond or second by second basis.

James Governor (37:54)
Yeah, yeah. I mean I I think that’s one of the the to your point

Or rather, the the the key thing is is is closing the loop. I mean, you know, control theory that that that observability is only valuable in so much as it allows us to take action to deal with the fact that again we are actually talking about customer behaviors, the purpose of the system. and I mean of course you have valued colleagues that that are building observability tooling. So, you know, throw them a bone, Marek.

Marek Poliks (38:05)
Exactly.

Yeah, absolutely.

So you asked me to throw my colleagues a bone and I let me do that. That’s OK. Yeah, let’s do that. Yeah, absolutely. And I think like what I’m really, really interested in is like to me, like a future of LaunchDarkly is is absolutely what is being enabled by the observability team right now. But what we’re ultimately building is we’re building essentially the kind of primitive or the control scheme

James Governor (38:35)
yeah. Let’s do that. Like throw your collar bone here.

Marek Poliks (38:58)
for software factories. And that does require observability in the sense that that requires the ability to continuously pull information from production and to enable that information to drive feedback loops. That is really, really critical. That’s exactly what you were referring to earlier when you invoked control theory, right? This idea that like the goal is to take information, especially, and that information is objectively more useful the closer it is to the moment in which it was actually captured, right?

And it’s objectively more useful the more that information contains, let’s say, some kind of like salient characteristics of the end customer experience. Both of those components need to go through a continuous feedback loop through which code is being continuously revised. So the individual observability components of, let’s say, a single feature, and this is again why feature flags are so useful, or let’s say a particular agent harness are extremely useful.

James Governor (39:37)
Mm-hmm.

Marek Poliks (39:55)
and that you have the ability to then connect and retroject that harness’s influence on the metrics that it was experiencing in time with its next iteration. The ultimate goal, especially using a system like agent control, that you have an agent, it’s live, you have live information around how that agent is performing from evaluators, from logs, from metrics that are being sourced from user experiences, and that information is continuously going into a reinforcement loop, right? It should be a reinforcement learning loop.

through which that agent harness is gradually and continuously evolving. That is really objectively useful. The consumer of those logs, the consumer of that observability information at the small scale tends to be the system itself, right? If I think about, let’s say a user who is right now coding at this very moment in Cloud Code and they’re looking at two diffs that like they absolutely can’t parse and they hit accept, but like that.

Like that larger scale picture is really, really difficult to deduce. That smaller scale picture, when you’re looking at individual contributions from a single harness to a series of experiences as expressed in statistical means and necessarily as part of a continuous live running system, that’s the future. Like that’s the future of LaunchDarkly. And so I feel like, you I don’t think it’s a contradiction at all to say that observability in the sense of passive observance of phenomenon is over.

And the new phase is about continuous feedback loops, a continuous, you know, sourcing of information from prod, from customer behavior, from, again, this idea of the purpose of the system being what it does into the continuous refinement of that object. And it becomes, you know, as you do that, it becomes more and more difficult to observe that system as a whole. And ultimately, you are going to ultimately attend to the behavior, to the outcome. That’s the thing that like, as you move up the management ladder, that’s ultimately the thing that you attend to. And that’s what Stafford Beer taught us back in like 1965.

James Governor (41:55)
Amazing. I love ending a podcast in 1965. I am all about I am all about theory and history and how we got here. I think that’s a great that’s a great rap. We we we threw threw them a bone. We ended up in nineteen sixty-five. Nineteen sixty-five we’re you know, I’m pretty sure we’ll you know we’ll we can get our flares out going into the late sixties. It’s gonna be it’s gonna be great.

Marek Poliks (41:57)
Hehehehehe

Great.

James Governor (42:23)
But yeah, no, that was fascinating conversation today. thank you very much for joining the MonkCast. It was super fun. I hope you’ll come back. I’m sure there will be more room for further spicy takes from Marek Poliks, head of AI at LaunchDarkly. so yeah, thanks very much for joining us. For all

Marek Poliks (42:49)
so much for having me this is really fun James we got to talk more often I really like hanging out with you for real

James Governor (42:53)
Well that is a hundred percent true. And i I mean it’s hundred percent true that it is fun hanging out with you. I’m not saying I’m amazing. But but you know, we can we can live with that. We can live with that. Occasionally we have to take the win. but if you enjoyed this show, you should do all the things. You should share it with your friends, you should get your mother to listen to this podcast, she will enjoy it a lot. share it, like, subscribe. but most of all, once again, thanks so much, Marek.

Marek Poliks (42:55)
you

Hahaha!

More in this series

Conversations (142)