Why do some teams see massive speedups from AI coding agents while others barely notice a difference? In this RedMonk New Builders Conversation with AWS, Kate Holterhoff talks with Liam Greenamyre, Principal PM of Developer Platforms at AWS, about the uneven impact of AI on software development and how to restructure work so agents can succeed. Liam explains why making tasks verifiable is the single biggest lever, then walks through the Agent Toolkit for AWS: a remote MCP server, 100+ skills, plugins, and a rules file that work with Claude Code, Codex, Kiro, and more. They dig into how a single RunScript tool covers 15,000+ AWS APIs token-efficiently, how IAM context keys separate human and agent actions, how agents discover new services like Aurora DSQL, and why isolated dev accounts make safe sandboxes for autonomous agents. Liam closes with a bold prediction: the biggest changes ahead are organizational, not technological.
This New Builders conversation is sponsored by AWS.
Links
Transcript
Kate Holterhoff (00:04)
Hello and welcome to this RedMonk New Builders Conversation with AWS. My name is Kate Holterhoff. I’m a senior analyst at RedMonk. And with me today I have Liam Greenamyre. He’s the principal PM of developer platforms at AWS. Liam, thanks so much for joining me today.
Liam Greenamyre (00:19)
Thank you so much for having me. I’m excited to talk to you today.
Kate Holterhoff (00:21)
me too. Okay. So we are going to be grappling with a lot of things that have been top of mind for a lot of developers that I know. namely what it’s like to write software i in the the AI era, what the SDLC is gonna be looking like six months in the future, a year down the road, all of these things. So let’s dig in and let’s start with introductions. Talk to me about what it is that you do at AWS.
Liam Greenamyre (00:45)
Yeah, so as you mentioned, I’m a principal product manager in developer platform. I focus specifically on what we have started calling our agentic developer experience. So really focused on how we can make customers using coding agents as effective as possible on AWS. And that consists of a couple of things. So making sure that agents have the right guidance and knowledge that they need, that they can
interact with AWS services very efficiently and in a very token and context efficient way. And they have the right guardrails so they don’t go off and do something crazy or destructive in your account. And so I work across a bunch of different teams within AWS trying to make sure that everything that we offer to customers is the best possible
experience not just for humans anymore but also for the coding agents that are increasingly becoming a way that customers interact with AWS. I’ve been in this role for only about six months. prior to that I spent five years in our billing and cost management org. So I don’t know if anybody listening is a FinOps practitioner or has experience with FinOps but that was my bread and butter for about five years. two of two of those years I was spent
primarily focused on our customer facing AI products in the FinOps space. so I brought a lot of that kind of domain expertise in in building AI systems. and then fun fact, I actually started my Amazon career in the retail side of the house. and so I was the category manager for the rugs, home decor, and artwork retail categories. And so I like to say I was negotiating about truckloads of
Yankee candles before Prime Day and then transitioned into building cloud infrastructure.
Kate Holterhoff (02:40)
God, I love that. I love that. Okay. Yeah, there’s two sides of the house there. That that’s exciting. Okay.
Liam Greenamyre (02:44)
Mm-hmm. Yeah, very different.
Kate Holterhoff (02:47)
All right. So so you and I have have talked about some of this, these broader issues around AI and development before. And so but I I’m hoping we can rehash some of it for this podcast. So talk to me about some of the unevenness that you have seen.
When
it comes to the adoption of AI by software development teams. Because, the sense that I get is that it’s not all successful. Like there there are some folks who are are struggling with this still. and so I I’m interested to hear some of the details there. So yeah, d without me putting more words in your mouth, yeah, how would you characterize that situation?
Liam Greenamyre (03:28)
Yeah, for sure. So I think when we look, you know, both across internal teams at Amazon, but also with different teams and customer organizations, what we see is that some teams and some organizations and some projects seem to get this massive speed up with AI, and some seem to get maybe marginal benefits, and some seem to get kind of almost no benefits at all.
But it’s not just at the organizational level, it’s also at the individual level. I think what a lot of us experience is, you know, whether it’s from first hand experience or things that we like read or see in the media, you see you know, you either experience or or hear about these kind of wild success stories where somebody used AI to do something that just sounds
you know, monumentally complex or would have taken months or years and teams of developers to do and they just let agents run with it and it produced a great result. so you hear about those. And then you also oftentimes experience these frustrating failures, right? Where you try to use AI for a development task and it’s just not really helping you. You’re churning, you’re spinning, it’s not really giving you the results that you want. And then
That’s kind of all interspersed with sort of mundane improvements, right? Like you implement AI in a workflow from maybe like code reviews or some sort of operational debugging, and you get, you know, some marginal improvements. and I think w what all of those things add up to, the fact that you can experience wild success and frustrating failures and mundane improvements all at the same time, I think that can be really disorienting for
both for individuals and for organizations. So for individuals, when you get a task or you take on a new project, it can be hard to guess like how much speed up am I am I gonna get with AI from this? Is this something where like, I can just like shoot off a prompt to Claude code and it’ll just go do it for me? Or is this something where AI is gonna actively slow me down? so it’s g it gets really hard for people to plan their own work.
But then at the organizational level, it becomes really disorienting too, because I think a lot of leaders are asking themselves, like, well, I’ve got project A and project B both are using AI really extensively. How come project A is moving so much faster than project B? Or if it’s so much faster to write code now, how come I can’t like ship my whole roadmap faster? How come I can’t expand the scope of problems that my team is taking on?
So that’s kind of what I what I mean when I when I say that it’s like it feels uneven with respect to the the impact of AI that that individuals and organizations are experiencing.
Kate Holterhoff (06:20)
Okay.
And so since you have seen this, I’m interested in what you’re telling these organizations ahead of time. Like how can they have the best chance of success and avoid these failures that you’re also seeing, or these slowdowns or, you know, problems with anticipating timelines, delivery dates that potentially don’t align with with the reality.
Liam Greenamyre (06:46)
Yeah. So I mean I think that the the best way to overcome some of this unevenness is just by understanding some tactics that you can use to kind of restructure the work in a way that is gonna make AI more effective at solving that problem for you. The way the way that I kinda like to think about it is you wanna set up the work in a way that
your agents can kind of bash their heads against the wall for you. and what I mean by that is you want to structure the work in a way where you can throw lots of agents, lots of tokens at the problem, knowing that you’re going to get good quality output at the end and that your agents aren’t going to make any sort of catastrophic mistakes or errors along the way. And so there’s a bunch of different kind of tactics that I can that I can talk about on this front, but the biggest one
that I would encourage listeners to take away is focus on how you can make your task verifiable. So what I mean by that is you need to figure out how you can structure the work in a way that agents can sort of check whether they’re doing the right thing and then get fast feedback on that. And so one of the types of
like wild success stories that we that we see and hear about is things like rewriting a whole system in Rust. And I think in you pi as you peel back the layers, one of the things that one of the reasons why those projects tend to be successful is that they’re successful when you already have a big test suite for the system that you’re rewriting. And what that does is it makes the task verifiable. The agents can rewrite a piece of the code, run a subset of the tests,
And evaluate whether it’s working. And then they can know, hey, if all these tests cases pass, then I’ve completed I’ve completed the work. and so if you can find a way to structure the work so that it is verifiable, you have a much higher chance of success than if you give your agents a task that is vague or open-ended or
Subjective with respect to what good looks like.
Kate Holterhoff (09:15)
That is so helpful and extremely topical. I mean, it seems like every other day in the news we’re hearing about some massive project that has devoted, you know, X number of tokens to migrating from one language to another, invariably to, improve the performance of of some
Liam Greenamyre (09:31)
Mm-hmm.
Kate Holterhoff (09:32)
application. and yeah, it it just seems like the barrier to these, huge migrations, which would have been unthinkable years ago, is just lowering and lowering. And so we’re just
it’s it’s I I I think a reality that a lot of organizations are still trying to come to terms with that, yes, this project that we’ve put off, all of all of this technical debt, things like that, we can suddenly approach it now. It’s it’s it’s not out of reach. It’s not it it makes business sense for us to look into to fixing this. So yeah, to so makes a lot of sense. and I I’d like to use that sort of
problem
statement as a jumping off point to to talking about something that you’ve been extremely involved in, which is the AWS Agent Toolkit, which does seem to address a lot of these problems that that y you’ve laid out for us here. So talk to me about what that is and what it accomplishes.
Liam Greenamyre (10:27)
Yeah, so the Agent Toolkit for AWS is really a suite of a couple of different things, all of which are designed to help your agent be more effective when it’s using AWS. So it works with any coding agent you’re using, so whether that’s Claude Code or Codex or Kiro or many of the open source harnesses that are out there. And it it consists of a few different parts. So one is the AWS MCP server.
So this is a remote managed MCP server that gives your agent access to any AWS service or API, which we have more than 15,000 of. So it’s very broad, broad surface area. And I can talk more about how we do that in a very token efficient way. and then the second piece that it’s that the toolkit consists of is skills.
So we’ve got a portfolio of over a hundred different skills that cover are designed to cover the sort of the breadth and depth of AWS. the third component is plugins, and plugins are relatively recently introduced format or protocol that bundle together skills and and MCP servers and other ancillary things like hooks.
Into a single install, so it makes it really easy to kind of distribute and install these things. And then the fourth component is a rules file. And a rules file is very simple. It’s a markdown file that you kind of stick in your agent’s context that gives it some some high-level instructions about how to use AWS and how to use the tools that are that are available to it. So those are the the things that the Agent Toolkit consists of. Now, how does it help with these problems?
There’s there’s a couple of main ways. So the first is giving your agents the right context. so one of the other major factors in whether your agent is going to be successful with a task is whether it has the right context, right? Whether it has the right sort of steering, the right access to knowledge, whether it can get the right facts and information and procedural knowledge at the right moment.
And so the Agent Toolkit helps with that not only by providing all of these skills, but also by giving your agent really efficient ways to, for example, search documentation and find whether it’s best practices or the API model for a particular service or troubleshooting guidance. All of that sort of wealth of information is available to agents as they as they go about their work. and so that makes them not only much more likely to
accomplish the task at hand, but also more efficient in in doing so. so that’s the first piece. And then the second piece where the Agent Toolkit helps solve some of these problems is helping prevent your agent from doing something that you don’t want it to do. so we I talked earlier about you want to let your agents sort of bash their heads against the wall for a on a problem for you. Now
One of the key aspects of that is if you’re gonna let your agents run for a long period of time and you’re not gonna be sitting there babysitting every tool call, you need to have a pretty high degree of confidence that it’s not gonna do something stupid, right? Like drop a production database or take down a service, right? So you need to make sure that it’s operating within a certain set of guardrails. And so that’s one of the big things that the AWS MCP server does is it provides a a set of guardrails that allow
You to define what permissions your agent has. Now, the key way that we do that is with IAM context keys. So you can most most developers that I talk to, they don’t set up, for example, a different account or a different identity for their agents. They give their agent access to their own AWS profile. Now, when your agent does that and uses, for example, the AWS CLI.
That means that anything that you can do, your agent can do. And so the same permission that you give to, for example, a senior engineer, you’ve actually also given to their Claude code, which may not be what you meant to do. Now, with the AWS MCP server, what you can do is you can define via these IAM context keys, you can define permissions such that even if
Your agent is using an admin profile, you can set restrictions on the actions that could be done via the MCP server. And you could, so for example, you could restrict it only to a set of non-destructive read-only actions, or you could choose to allow only a low risk set of read only of mutating actions. So you can kind of set those permissions in a way that
Creates the proper set of guardrails without having to give your agent full access to everything that you can do. The other nice thing about these IAM context keys is that it allows you to audit what agents are doing and to distinguish between human actions and agent actions. So if you have Claude Code or Kiro, for example, just using the AWS CLI.
and then an incident happens, you’re going back and looking through the logs. You you can’t tell like was this a human, was this an agent operating autonomously? Was there a human in the loop on this? versus when you use the AWS MCP server, you get both cloud trail logs and cloud watch metrics that distinguish between human and agent identities. And so you can actually see who is the
The human user, as well as the fact that this was an action that came via the MCP server, and so you know that it was an agent that did that. So, in summary, those are kind of the two main ways that the Agent Toolkit helps to helps to address some of these challenges, is that it gives your agents the right context that they need, and it also establishes a set of guardrails to make sure that your agents don’t go off and do anything stupid while they’re operating autonomously.
Kate Holterhoff (17:00)
Okay.
Yeah. That that helps me to understand. And also I think what you know what I’m hearing from folks is that they they absolutely need these guardrails right now because, you know, not
Liam Greenamyre (17:13)
Mm-hmm.
Kate Holterhoff (17:13)
only do we have, the the permissions that you know could get out of hand, but also that what it is that developers are doing on a day to day basis is just shifting so dramatically. So having the ability to use their personal
harness in order to access AWS services in in a the best way seems absolutely essential. Like, making sure they they have the most up-to-date documentation, that they’re not maybe looking at that that the agent isn’t looking at some blog posts that, is out of date or getting their information from a source that maybe just doesn’t have
the best way of going about a certain intended outcome.
Liam Greenamyre (17:53)
‘Cause the thing is like the the
agents like they don’t know what they don’t know. and
Kate Holterhoff (17:59)
Well put.
Liam Greenamyre (17:59)
oftentimes they will you know, when you ask them to do something, they’ll sort of run off and go and go start executing without, you know, grounding themselves in do I have the information that I need to to go do this. And I think one area that this is particularly that this can be particularly painful is when AWS has
Release new services, which obviously we
Kate Holterhoff (18:24)
Mm.
Liam Greenamyre (18:24)
do all the time. one of the ways that I heard an engineer put it to me that I really liked is we have tens of thousands of people whose job it is to change AWS every day by making it better for customers. Right. And so what we see is that when AWS releases these new features, it can take a long time for models to become aware of.
those things. So whether it’s Bedrock AgentCore or DSQL or S3 vectors or Lambda MicroVMs, agents r really don’t know about those those features for a long time unless they are either you know searching the web or they’re using the documentation. And what we see is that that can result in like pretty poor architecture.
Decisions in some cases. Like I’ve seen a lot of examples where, you know, an agent, you know, a customer has a use case that’s a perfect fit for a new feature that we just released, whether it’s agent core or it’s you know, vector search that S3 vectors would be a great, a great fit for. and agents don’t know about those features unless they have the Agent Toolkit installed. And so what you see is that they
you know, try to stitch it together from multiple different services, or they use the wrong compute primitive, or they like hand roll some of the functionality that that these services offer out of the box. And so you can end up with subpar architectural decisions just because agents don’t know about the latest offerings from AWS.
Kate Holterhoff (20:10)
That’s
huge. And and frankly, I’ve I’ve gone through this myself where I was doing more of the ClickOps style, but yeah, the sort of advice that I was getting from the agents didn’t align with best practices. So it’s great that now we have a solution for this that is gonna help me to have access to the information that I need in order to, yeah, be more successful and and maybe avoid the ClickOps, you know, be able to automate some of this. I mean, this seems extremely valuable.
Liam Greenamyre (20:36)
Yeah. And what we’re hearing from what I hear from customers using the Agent Toolkit is oftentimes they tell me like, hey, a bunch of things that I used to go into the console for, I’m just doing via my agent. whether it’s, you know, trying to understand why a CloudWatch alarm is is firing or running logs queries, or you know, even you know checking costs and understanding why costs increased or decreased week over week. These are all things that, you know, in my
my daily life I tend to turn to agents for and I’m hearing from more and more customers that this is, you know, the main way now that they interact with AWS.
Kate Holterhoff (21:16)
And that’s huge. I mean, we hear more and more about folks making their own dashboards to you know, be able to monitor things like cost. Is that one of the use cases? And I’m thinking of just developers, but it sounds like there’s also the advantage of being able to to monitor costs in a in a format that makes the most sense for that end user.
Liam Greenamyre (21:34)
Yeah, for sure. I mean AI is is definitely making a big impact in the in the realm of FinOps. you know, when I was when I was in that space up until about six months ago, we talked a lot about both AI for FinOps as well as FinOps for AI. So both using AI to understand your overall cloud spend as well as applying the practices of FinOps to this exploding
subcomponent of of your AI spend. And so it’s definitely making a huge difference there. I think that’s wr the reason why customers are not only using you know the Agent Toolkit to help monitor monitor and understand and control their costs, but also we’re using more sort of out-of-the-box offerings like the FinOps agent. And what we’re we’re really seeing like a lot of enthusiasm from the FinOps community in in using
The FinOps agent in general or sorry, FinOps agent specifically, but also just AI in general to help accelerate all of the analysis and reporting and finding optimization opportunities and sharing those with with engineering teams. There’s a lot of routine work that AI can help to automate there.
Kate Holterhoff (22:51)
Yeah, yeah. I’m interested in the genesis of the toolkit because I think what’s unique about this service is that it it stitches together MCP plugins and skills and I I I tend to think of them as being different from a lot of other, you know, vendors. So were they separated up until now or like, how d how did this come to be?
Liam Greenamyre (23:14)
Yeah, for sure. So w the thing that it really started with was MCP servers. and that was honestly the that was the protocol that was released first and that caught on really quickly. And so we had we had a a a number of sort of prototype MCP servers that were released. We consolidated that into the AWS MCP server.
Which is also a remote MCP server. So it’s a little bit more sort of production ready. It’s not just, you know, code that’s running locally on your machine where you have to install it and install dependencies and things like that. So it’s a little bit more production ready. and you know, the benef the other benefit of the AWS MCP server is that it, you know, consolidates the functionality to to interact with any AWS service and API, like we talked about. So that was the thing that we sort of had.
First, we launched that in preview in November, December of twenty twenty-five, so roughly a decade ago in AI time. and
We
were working towards a general availability launch of that and what what happened along the way is that the world changed, right? So skills became a very prevalent way that people were providing additional context to their agents and and helping improve task success rates. plugins came out as well, again, sort of a way to bundle together different different aspects of agent tooling into a single install. And
So what we realized is that we needed to be A, meeting developers where they were, and and B, we needed to make sure that we were not narrowly focused on like one aspect of the agent tooling landscape. and so what we decided to do is to think a little bit bigger and think a little bit broader, and that’s where the Agent Toolkit came in.
and we we decided to do this in a way that we would sort of bundle together the skills, the MCP servers, the plugins, the rules file, and help customers understand this is kind of all one thing. It’s one offering that is helping you improve your agent performance and success rates and efficiency on AWS. but that you can also sort of mix and match the components, right? So you could install just the MCP server or just the
Just the skills depending on what your needs are. but again, it’s all kind of in this in this area of improving agent performance on on AWS. And so that’s where the toolkit name came from, and that’s how the overall toolkit came about. And so we launched this in May of this year, and we have seen
really really rapid adoption and growth amongst AWS customers. As I mentioned, it’s really becoming like a primary way that a lot of customers are interacting with AWS, and hearing, you know, really positive customer feedback and asks for for a lot more for us to for us to do to to help help agents be successful.
Kate Holterhoff (26:45)
That’s awesome. and okay, so I have so many questions after that. But l let me begin with the the skepticism that some folks have about having heavy skills and heavy MCP or whatever. Because when I try to wrap my head around all of AWS, I’m like, you’re telling me all of that is crammed into this MCP. Like, I don’t know if I want to be dragging that around with me just to do like a an S3 thing. You know, like this this seems like maybe
A little too too much. so talk to me about about that. Like what how do you ensure that this isn’t just this like heavy behemoth that’s using up all of your tokens for for every small little query?
Liam Greenamyre (27:27)
Yeah, for sure. So we’ve had a pretty sort of maniacal focus on token efficiency since we started this work, which has required us to think differently about how we’re gonna support this breadth of AWS services. So when MCP servers started, the the typical way that it would work is you would sort of wrap the API. So let’s say you’ve got a basic CRUD API that you wanna want your agent to be able to interact with. Well
Initially, what you would do is you would sort of write a tool spec that translates all of those API actions into things that the agent can interact with by giving like descriptions and describing the parameters that are required, and then the agent would invoke that tool, and then the tool is basically just proxying that to to that API. But obviously, that is not gonna scale.
to the sort of 15,000 APIs that that we that we’re talking about here. And so what the AWS MCP server does differently is it exposes actually a single tool that gives agents access to the entire portfolio of AWS APIs. And the way that we do that is through a tool called RunScript, which sounds very very simple but powerful, right?
so basically what it allows agents to do is essentially write Bodo 3 code, so our Python SDK that can invoke AWS APIs. And I think the key insight here is separating the interactivity from the instructions about how to use those APIs. So the run script tool gives agents the way that they can interact with any of these.
services, but it’s actually the skills and the and the documentation search and these sort of knowledge tools that give the agents the understanding of like what is the you know what are the parameters that I need to use or how do I use this API. And that’s this ends up being much more efficient, token efficient in a couple of ways. One, it’s important to realize agents already know how to use a lot of AWS. So you your agent doesn’t need
a skill, for example, to tell it like how to describe your EC2 instances. It knows how to make that. It knows how to make that that API call. And so in that case, just allowing it to sort of use Bodo 3 to make that call is much more effective than taking that API and wrapping it in like a tool description just for that API and giving that to the agent. and then
The other thing that’s really token efficient about this approach is it allows agents to compose different API calls into a single tool invocation and then also parse the results before that gets passed back into context. So let me let me maybe give a concrete example that that helps to helps to illustrate.
So let’s say you’ve got like a thousand of a resource type. So you’ve got a thousand S3 buckets, and then you need to, for each of those buckets, you know, check one aspect of configuration, and that comes through another API. So with the sort of old approach or the classical approach to MCP, your agent would need to
List the buckets and then for each one of those make a different tool call to get the detail that it needs about it. And so that’s massively inefficient and uses a ton of context. Now, what you can do with this run script approach is an agent can basically just write a script that does all of that server side, parses the results, and so it only gets back A, it can make all of those API calls in parallel.
but then B it can parse the results so that it doesn’t have to for example get the full list and and all the descriptions of each. It can just get the fields that it needs, filter server side, and then only pass back the results that it’s interested in. And so this is a much, much more token efficient way to do complex operations on AWS.
Kate Holterhoff (32:17)
super helpful.
Liam Greenamyre (32:17)
Yeah.
Kate Holterhoff (32:18)
What I want to hear about too is something that you mentioned about discoverability. So a
Liam Greenamyre (32:24)
Mm-hmm.
Kate Holterhoff (32:24)
thing we at RedMonk have been talking to a lot of folks about is this idea of like Agent Experience (AX) and then the related issues of like, how is it that users are going to be finding
Services. These dev tool vendors are suddenly in this new world where SEO doesn’t matter in the same way that it it did in the past. And so I I’m I’m curious you know, and when I think of AWS specifically, I think of the certification programs, which have been such an important part of how like cloud engineers have learned about all the AWS products. And I kind of it it almost feels like there is a part of this new reality.
w where we’re using our agents now to to do the work that these certifications have done and that platform engineers have had to to rely on, using search engines to to to deal with, or maybe attending conferences so so just everything has changed around discoverability. And discoverability is this unique challenge at AWS. So have you framed it?
in that way when you’ve had these conversations in terms of like, ensuring that the updates to the toolkit are what, weekly, daily, I don’t know, to to make sure that you do have that discoverability all the time. And and like how are you making sure that that these users are actually going to be able to find what they need and and maybe even be introduced to solutions that they hadn’t thought of in the past?
Liam Greenamyre (33:51)
Yeah, absolutely. So in terms of the pace of updates, we release updates particularly to the skills portfolio on at least a weekly basis, usually more so
Kate Holterhoff (33:59)
Okay.
Liam Greenamyre (34:00)
a little bit more like daily. and you know, obviously that’s that’s the pace that’s required when when we have so many people working to make AWS better every day and we release so many new new products and services and and features. but I do think that actually
agents can be really helpful with that sort of discoverability problem because if you can if we can provide the right context to agents via these skills, what we see is that agents will naturally discover the right service for their use case without having to, you know, go like be educated about it, for example, or you know, read a blog post or
you know, get targeted with an ad, for example. And so it can provide a really natural discovery point where if we you know, if we provide the right the right skills, then an agent can say, Okay, I have a, you know, a data intensive application. I need to choose what the right database service is, let me go check this database skill.
I’ll just you know, I’ll it’s got a description of like which service I should use for which use case. And hey, there’s this thing that I haven’t heard of called DSQL, but it sounds like it’s a great fit. And look, there’s a another skill just for Aurora DSQL where can learn all about how to use it. So it provides this sort of natural progression from what’s your use case to what service do you need to how to go build with it.
that is I think it just provides a really natural flow for discoverability in a way that’s g that can be harder to do with when you’re sort of trying to get attention and share of mind from a human developer.
Kate Holterhoff (35:56)
Yeah.
So you have mentioned the IAM story and how permissions are such a big part of the toolkit. but I I want to hear more about that because the danger that so many folks worry about is like deleting prod, you know, like losing a database. yeah, horrible, horrible things. And I can’t think of a more
critical service than the one that AWS is providing, right? You know, this is not just like some, superficial update. This is where everything lives. And
Liam Greenamyre (36:29)
Mm.
Kate Holterhoff (36:30)
so I just I I guess I just want to hear a little bit more about, what’s the hip term now? Guardrails, right? and how you’re ensuring that these terrible things aren’t going to happen. Like, is there a rollback situation?
I guess yeah, ca can you just talk to me more about how you are addressing this very relevant fear? especially as I, we haven’t even really framed it in terms of personas, but I mean I get the impression now that, more folks are gonna have access to some of these, again, critical services that maybe hadn’t in the past, like PMs, maybe folks who historically have identified as like front-end engineers or designers, like suddenly they are now.
able to do a lot of this this work in a hands-on way. So anyways, d yeah, talk to me about guardrails, you know, ha ha help help me to to mitigate these fears.
Liam Greenamyre (37:19)
Yeah, for sure. So I’ll I’ll come at that maybe from a few different angles. So, you know, I think the the AWS MCP server gives one form of guardrail, which is these IAM context keys. And that again allows you to kind of distinguish between what it what can a human do versus what can an agent do. And so I think that’s one layer. But there’s actually in practice, there’s it’s you can think of it as sort of defense and depths, like we’ve always sort of thought about these things for for security.
There’s multiple layers of of guardrails and controls that you might consider. and so one is what are the permissions that you do or don’t allow your agent to do just within like on your laptop with the on the client side, right? And so that in practice could mean things like what are the set of commands that you allow Claude code to run?
and so a simple thing that you can do, for example, is you could say, Claude Code, you are not allowed to invoke the AWS CLI directly. You’re only allowed to go through the MCP server. So that in and of itself is is something that you is a form of guardrail that you can that you can put in. having the right IAM permissions and roles is a is a key piece of it. but beyond that.
There’s a lot that you can that you can do to introduce additional types of guardrails. So one, for example, is running your agents on a an isolated machine, right? So it’s not just taking down you know production systems or or making a mistake with your AWS infrastructure that that folks worry about. It’s also like
For example, a bad RMRF command on your laptop that bricks your machine, right? But if you’ve got your agent running, for example, on a Lambda MicroVM or you know, even a classic EC2 instance, then you’re less worried about that, right? Because you can just reboot it and and and keep the work going. So putting your agent in a sandbox is definitely a big piece of that. but then when it comes to AWS, what I hear some of our more sort of
mature customers doing is things like, hey, you know, I’m gonna I’m I’m working on the architecture for a new system or I’m building building a new feature or a new system that involves some infrastructure. I’m gonna create a brand new, fresh, isolated dev account just for my agent. And I can let it it you know coming back to the phrase, I can I can let it bash its head against the wall.
in that account knowing that it doesn’t have access to anything that matters. And so it can iterate, it can do the ClickOps, ClickOps, quote unquote. You know, it can it can execute the commands directly to to to build the service. It can iterate as much as it needs, run tests. And then when it’s done, when it’s ready to actually put something into production or or you know promote it beyond this dev account, you have it then go write a pull request
For with IAC that will actually go through human review. So you this is actually a good example of kind of structuring the work in a way that your agent can bash its head against the wall. You put it in the sandbox of an isolated dev account, and it can go do whatever it needs to do to create the service and iterate as much as it needs. and then the you know, when it’s actually ready to to promote, then it actually goes through human review by submitting an IAC code change.
Kate Holterhoff (41:03)
Okay,
so we are we’re winding down here, but I I wanna I wanna hear some of your thoughts on where all of this is going. so you’re at this privileged position. AWS is doing a lot of really cutting-edge research on how it is that agents are changing the SDLC and and infrastructure and how we’re all going to be sort of encountering
the acceleration that this enables, the these new security issues that didn’t exist. So yeah, all of that is to say, what I want to hear about is your thoughts on like a roadmap, like future looking ideas. Like where would you say that AI software development is is heading in 2027, and maybe, you know, maybe even I I’d ask you to go a little further than that, like, you know, five years from now. What what do you think that’s gonna look like?
Liam Greenamyre (41:54)
Yeah, I mean this is an area where you know predictions about the future tend to have an unusually short shelf life.
Kate Holterhoff (42:01)
Yeah.
Liam Greenamyre (42:02)
but you know one of the things that’s that’s interesting is I think that a lot of I think a lot of the big changes are gonna be more organizational than technological.
Kate Holterhoff (42:16)
Mm.
Liam Greenamyre (42:17)
I think that even if
Even if models stopped improving today, or even if harnesses stopped improving today, I think there would be years of work ahead of us to really figure out how do you how do you organize software development teams w in a world where writing the code is no longer a primary bottleneck of of actually shipping value to customers. so
you know, how do you organize these are things that we’re these are questions that we’re asking ourselves internally, right? Like how do you or organize teams and organizations and where do you put your focus? How do you for example, streamline decision making when it’s so it’s so fast to actually build these systems sometimes that decision making about what’s the right thing to build can become a bigger bottleneck than actually than actually building the thing. and
You know, beyond just kind of org structure types of questions, it’s also, you know, like I said, about how how people structure their own work. I think people will spend I think developers will spend a lot more time thinking about architecture than thinking about syntax. I think that developers will spend a lot more time thinking about how to verify output rather than specifying it.
the specific steps that are going to be taken to generate that output. And I think that as these changes continue to ripple throughout the the SDLC, managers and and and developers will spend a lot more time thinking about where the bottlenecks are moving to when, you know, when again, when software, when writing the code,
becomes so much faster, you see the bottlenecks move, move elsewhere, whether it’s your build, whether it’s your code reviews, whether it’s your testing, whether it’s your deployment, or like I said, upstream with with decision making. I think that all of that all of that work of really rethinking how we do software development, in my mind, that’s almost more interesting and and more where more work needs to be done than just continuing to improve the technology.
Kate Holterhoff (44:43)
Yeah, yeah, yeah. I love that answer. Yeah. Thinking through the the human problems, the fact that this is something that’s going to be yeah, or organizational wide, that this is really about structuring the people behind it. Love it. Okay. Talk to me about where folks can learn more about the Agent Toolkit. Where do you direct them?
Liam Greenamyre (45:00)
Yeah, so you can just search for Agent Toolkit for AWS and you’ll find our GitHub as well as the product detail page. And then if you’re already an AWS developer, you can use the AWS CLI and just type AWS configure agent-toolkit and it will walk you through installation and setup.
Kate Holterhoff (45:19)
Amazing. And for folks who want to hear more from you, Liam, personally, do you have a social media presence? Are you on LinkedIn? Like how can folks keep in touch with you?
Liam Greenamyre (45:28)
Yeah, I’m I’m on LinkedIn and you can reach out to me there.
Kate Holterhoff (45:31)
fantastic. I have really enjoyed speaking with you today, Liam. Again, my name is Kate Holterhoff. I’m a senior analyst at RedMonk. If you enjoyed this conversation, please like, subscribe, and review the Monkcast on your podcast platform of choice. If you’re watching us on RedMonk’s YouTube channel, please like, subscribe, and engage with us in the comments.





