Engineering Teams in the Agentic Era

Nate Umlauf Mar 26, 2026 30:00 64 transcript lines 18 terms defined Watch on YouTube Source page

Learn how to restructure your team's workflows around foreground and background AI agents. We'll cover which tasks to delegate to fast vs. slow agents, how to rethink meetings and sprint cadences, and practical frameworks for rolling out async development practices across your organization.

Terms in this video

Transcript

Cool. Well, it looks like we've got just almost 200 people. So, we can [clears throat] go ahead and and maybe dive into to the content here. Uh before we get into it, I'll just give kind of a quick intro. Uh my name is Nate and [clears throat] I'm on the AI deployment team here at Cursor. And so, really what that means >> [clears throat] >> is I just work with a lot of our large enterprise customers and help them roll out Cursor successfully. So, it involves a lot of technical enablement and kind of just

helping folks think about deploying AI at scale. And then, I also have Audrey from my team here. Audrey, if you want to give a quick intro as well. Sure. Hi everyone, nice to meet you. My name is Audrey. I'm also on the AI deployments team like Nate, based out of San Francisco.

Cool. And for today's topic, we're going to be talking about engineering in the agentic era. And so, you know, there's been kind of a lot of changes over the past year in terms of uh how everything works in terms of building product and software engineering. Uh if we were having this webinar a year ago, we might have called this engineering in the AI-assisted era, but now we're seeing, you know, AI agents are taking a lot of work off of engineers' plates. And so, that's sort

of what we'll be covering today. So, as a quick agenda, we'll start with just some of the trends that we've seen across all of our customers who are utilizing Cursor and just industry trends. Then, we'll talk about like what's actually changing about the job. We'll dig into uh sort of a a key topic of using like foreground agents versus background agents and how you want to split up the work between them.

And then, we'll take a look at uh some of the internal use cases that the Cursor team has seen really good success with these different agentic features. And then, we'll kind of close with uh how you can utilize Cursor to adopt some of these changes and best practices. Um as [clears throat] I said, uh feel free to throw questions into the the sort of Q&A section of the Zoom and uh we'll definitely have time to answer those at the end, but we'll also be taking

questions throughout the presentation. So, before [clears throat] we dive in, I just want to give a quick sort of lay the land on where you can use Cursor today. So, we are absolutely not just a like desktop application anymore. Um a year ago, that may more have been a little bit true, but now, you know, we are seeing folks utilizing Cursor's CLI that you can pipe into you know, any any workflow. We also have a partnership with JetBrains. And so, you can utilize the Cursor agent across IntelliJ,

PyCharm, and the sort of the full family of the JetBrains products. We also have the concept of cloud agents, which let you trigger and manage the Cursor agent from web, from your phone. Uh you can kick off agents via integrations we have with things like Slack or just via the API. And this means that, you know, you don't need to be at your laptop to uh use the the Cursor agent.

We also have uh a product called BugBot that does automated code review at the Oregon and repo level. And so, kind of the idea here is that Cursor surface area has expanded to essentially meet you wherever you and your engineers work across the entirety of the SDLC. Then, getting into [clears throat] a little bit of like the topic uh for today of how we've seen AI affect engineering orgs uh across, you know, all the different customers that uh use

Cursor, I'd say there's kind of a couple different things to to call out. The first and maybe uh one of the most interesting changes is that we've seen profiles of engineers go from what you might have called like T-shaped, where maybe engineers are really good at going deep in one or two subject areas, to, you know, expanding that range. And maybe engineers now have uh a lot of proficiency in going deep in a whole bunch of different disciplines. And because of this, we've seen

the number of people it takes to ship a feature end-to-end uh really reduce. So, instead of having an entire engineering team working on a specific feature or uh new product, you could see maybe a pod of two to three engineers. Or, in Cursor's case, we actually have single engineers that are owning entire features end-to-end. Uh and this obviously has, you know, increased the productivity of the org as a whole. Uh and it's also changed how the team

communicates. So, things like sprints and stand-ups are a little bit outdated for a lot of these engineering orgs and they've sort of transitioned into communicating in in different ways amongst product, engineering, and design. Then, in [clears throat] terms of kind of how we see this affecting like productivity, um again, it's sort of like that cross-functional product team uh or pod is replacing teams that maybe have really specialized silos. Uh and we see engineers working across

the full stack. Nowadays, you don't have to have a perfect understanding of a specific coding language in order to contribute uh on something that requires that language. You can learn it really quickly with AI. And we've also seen that this has dramatically reduced onboarding time for new engineers, where people are contributing in a matter of days or weeks as opposed to having some drawn-out month-long onboarding process.

I think one thing to kind of highlight here is that it's [clears throat] really important to note that we sort of think about this as not doing more with less developers, but essentially being able to ship a lot quicker and produce more value for the customers or whoever you're building for with the same amount of people. Uh and so, it's it's not, you know, a conversation of like reducing head count or or things like that and more of being able to just do more and and move quicker.

>> [clears throat] >> In terms of kind of where we see a lot of different orgs like on this uh I guess what we're calling like a maturity curve for AI, um we, you know, obviously started in the early innings where people were just coding manually and then that turned into uh at Cursor, we had Tab, which was more or less auto-completing your sentences and helping you write code quicker. And now, we've moved to a place where we are seeing AI used almost as like a

teammate or a junior engineer. We're able to delegate like an entire set of tasks or maybe uh really detailed tickets from like Jira or Linear to uh an AI agent, have them go work on that, and then come back to you with those results. So, we continue to see, you know, moving further to the right over here, where in the future we think that you'll be able to essentially have AI as operating as an entire engineering team or being able to sort of orchestrate multiple agents

at once that are doing a whole bunch of different complex tasks. And then, eventually maybe getting into using AI as almost like a co-founder or a builder, where you give more so the end state of what you want something to look like or a product, and then the large language model is working backwards from that to to engineer it.

Quick. Um just cuz it's kind of a good time for that question. Someone asked, "Has our industry come up with any good patterns to validate the production readiness of the rapidly increasing volume of LLM-generated code?" Yes. That that is a great question. Yeah, I mean, I think and we'll get into this a little bit more at the end as well, but I think from a Cursor perspective, one of the like biggest bottlenecks or maybe the biggest bottleneck now that we see is

not that like engineers don't have access to the right tools or that they can't generate code quick enough. It's more that like code quality is kind of becoming the bottleneck for a lot of these organizations, where we're developing, you know, code at 10x, 100x the rate that we previously were. And now, we just need to ensure that all of the same code quality metrics are staying consistent as you implement these new tools. Um there's definitely a

bunch of different ways that you can do this with in within an organization like at Cursor. For instance, we're using our uh code quality review tool BugBot to actually catch uh bugs before we ship things to production. And so, that's cut down on a lot of sort of back and forth with AI-generated code. We also are doing things more so in uh like before that stage, where all of our engineers are pretty consistently running what's called like a de-slop command uh that is

essentially taking a look at code and removing any patterns of code that are just like AI-generated slop more or less. Uh if you're interested in taking a look at that, that actually lives on our public marketplace. So, you can take a look and see at some of like the skills and things that we've baked into the code generation process that help with that. And then, we also have BugBot and a few other things that you could set up with like automations and

different cloud agents for reviewing code quality before that even gets to uh a PR or something like your your Git repo. And then, oh yeah, I see another question just around what does barrel-shaped mean in that previous slide context. And that just means uh going a little bit wider with like your breadth of skills as opposed to to being a little bit more narrow uh but deep in certain areas.

And then, here I just pulled some uh statistics around what we're seeing happening internally at Cursor and also amongst some of our customers. So, the way that we kind of think about Cursor now is we're less about helping you write every single line of code and more about building uh like a factory of agents that can help you produce software and product. So, you're more or less directing these agents as like teammates as opposed to I don't know, like steering the AI in

every single keystroke or every line of code that you're writing. And I think these three stats do a really good job of illustrating that. So, now almost 40% of PRs merged at Cursor itself are created by cloud agents. And this doesn't mean, you know, like edited by an agent. This is like end-to-end.

The cloud agent kicked off, wrote all the code, uh came back with some artifacts, and then that was pushed to to production. Uh we've also seen across our customers 15x growth in agent over the past year. And something that's really interesting is we're now seeing twice as many agent users as tab users.

And tab, the auto-complete feature I was talking about earlier, that was sort of what put Cursor on the map. And, you know, we've pretty much moved on from that into like a full uh agent-focused world. And so, I think this is, you know, these are some good stats of just showing like where this industry is going.

And then, >> [clears throat] >> if anybody on the call has uh checked out the Cursor blog, you might have seen that our CEO, Michael, wrote a blog post where he talked about what he refers to as like the third era of AI coding. And [snorts] these are kind of, I guess, the three eras you could talk about uh up here on the screen. So, first was starting with that tab functionality within Cursor, more or less using AI to auto-complete your sentences and write code faster. Then, we moved into

the second era, which was focused on like synchronous agents. So, having an agent riding shotgun with you while you're writing code. Um and I think what's really interesting about this is this agent this era probably didn't even last like a full year. Uh models became intelligent enough to do this maybe like 9 or 10 months ago where we were seeing people spin up an agent and like one-shot uh a prompt and the agent would get it right. And now we've almost seen a new

paradigm shift where uh people are kicking things off async and uh these agents are running in the cloud on their own VMs and then coming back with artifacts for the human to review. And so, yeah, I think each sort of era here is like marked uh by where developers are spending their like time and attention. So, with tab, it was like keystrokes and looking at the code. Then, in the synchronous agent area, you're kind of steering the agent. And now it's almost like reviewing or managing

an AI agent as a teammate. Um so, you could think about, you know, an engineer, maybe they have all these agents that are sort of like working under them as junior engineers, and then they're giving feedback and managing all those tasks in parallel.

And then, here's just [clears throat] kind of like a mental model of how you can think about the different ways of working with agents. So, we talked about uh kind of most of the ones on the left-hand side of like code-centric where you're looking at the code or agent-first where the agent's helping you edit the code. And sort of what we're focused on today in this webinar is these longer-running async agents. So, again, kicking off something in the cloud, having it run for hours or maybe in some

cases like days or even weeks. And then, the final piece is uh incorporating these agents across all the different tools in your stack and implementing them at different points in the software development life cycle. So, some of the things that Cursor has seen a lot of success with is using cloud agents in conjunction with uh tools like PagerDuty where maybe we spawn a cloud agent every single time there's an incident that's reported, and then a cloud agent is using the DataDog MCP to fetch logs and figure out what's

going on with an incident and then provide a summary to the engineer um as they're, you know, getting alerted of the incident as well. So, these uh this shift is kind of going from I guess what we would call like foreground agents where you're watching the agent write code and giving it uh feedback in real time to background agents where the agents are handling more of these tasks end-to-end.

And then, here's just a quick graph of like how Cursor has seen our cloud agent adoption spike. And this uh image is even a little bit old as, like I said, we're now seeing closer to 40% of our internal merged PRs being created by cloud agents. But you can see there was a pretty big uptick in adoption here once we released the concept of artifacts. And so, artifacts are basically we gave these cloud agents the ability to navigate their own machine. So, they're maybe making a front-end change, they go ahead

and spin up a browser and actually test those changes, and pretty much create a video of that, and then they link that video recording to a message to the developer and say, "Hey, here are the changes that I made, and here's an artifact to prove it." And so, that's kind of uh what I mean when I say that these developers now are reviewing the AI's code almost as like a junior engineer. Uh and at Cursor, this sort of concept of artifacts was a really big

unlock for our team adopting this this style a bit more. And then, this is just kind of reiterating that Cursor is not a platform to just write code, right? So, our agents are reviewing, testing, fixing, and deploying code. And this quality loop is is built into to the platform. And where we see this going over the next 6 months, uh and we've touched on a couple of these topics here, but, you know, what we've been talking about is a lot of this is moving

async. So, agents are going to the cloud. Um these agents can handle increasingly long tasks, and uh this sort of like cloud development uh in a lot of cases is going to be much faster than doing this locally. And then, we also are putting a lot of emphasis on the importance of like code quality metrics. And, like I said earlier, code quality has become a bottleneck for a lot of uh companies that we work with. And so, it's becoming like increasingly important to make it

easier within Cursor to ensure that you're reviewing code efficiently and not having, you know, any of those issues before these things get to production. And so, that's something we're really focused on from uh a product and like organizational standpoint.

And then, the last piece, and this may be more directly related to Cursor, is that uh we're going to continue providing enterprise controls for rolling these types of things out uh at scale. So, big organizations need a lot of visibility into, you know, who is launching cloud agents, where they're running. Uh there's a whole bunch of other like infrastructure uh things to think about as well. And uh we want to give you control over all these different deployments and all the other tools that that might live in

the in the stack of your SDLC. And so, that's something that our enterprise team is really focused on, and you can kind of continue to see uh things that we'll be shipping from a a governance and control standpoint. And Audrey, any questions that I should take live, you think?

Yes, I think there's there's a ton of questions around um thinking about how we can think about productivity in this new age of AI, how we can think about those new bottlenecks like product management, QA, and DevOps. Um I think a a lot of the answers um are really through how can we generate more automations and take advantage of uh increasing services like agent review and and BugBot, but curious if you have any other additional thoughts there. Yeah. Yeah, and that's a great question.

I mean, I think like internally at Cursor, we've definitely spent a lot of time thinking about this. And like one, I guess, interesting maybe like story or or tidbit is that automations is actually something that arose because our engineering team built like an early version of automations internally where uh all of our engineers really don't like getting paged at, you know, 3:00 in the morning or in the middle of the night and having to wake up and comb through logs to figure out what's going

on for an incident. And so, I think one of the earliest adoptions of the early version of automations was specifically for uh like PagerDuty uh incidents to trigger uh an automation. And then, again, that's that sort of we're having an agent look through DataDog and figure out like sort of what's going on there.

And so, I think a lot of this is like the tools are available for engineers to, you know, apply them to these other like disciplines, whether it's product management, QA, and DevOps. Uh but it's just very early innings in sort of where we are with that. And so, I think primarily people have been focused on like generating code, but I would imagine over the next like 3 to 6 months, you'll see a whole bunch of more like productized things around using

Cursor for these specific use cases. Um we've seen I mean, for product management specifically, uh we talk to a ton of product teams that are getting a lot of value out of Cursor for prototyping is maybe like the most obvious use case, but also, uh you know, doing things like customer research or combing through like interview transcripts and stuff like that. Um so, I think, yeah, our team is definitely thinking about how we can speed up the other parts of the process. And uh you can definitely be sure that we'll have

uh probably some more productized things in those other areas uh in the next you know few quarters here. Cool. And then I saw somebody asked about uh the the slide deck. Yes, we will send this out to everybody after the call. Um Cool. Then this is just kind of a really quick like mental model for thinking about when you want to use like a foreground agent or when you would want to use what we're referring to as like a background agent. So high-level like this foreground work is

for things that are interactive. So if it's ever a bit ambiguous or you want to steer the agent um you could think about if you're doing front-end work and you want to test those changes and get like real-time feedback and make updates or if you're doing something where you need to monitor like spikes or anything that might be really security sensitive, that's when you're going to want to use like this foreground agent where you can see what the agent's doing and sort of

monitor every single step in the process. For background work this is where the agent is just going to run on its own. And you can use background agents when you have like really clear goals and acceptance criteria. And so a lot of times here the output isn't like diffs or code file changes. Maybe it's a an image or a recording or logs that uh you can use those artifacts to kind of give a thumbs up or thumbs down once the agent has finished running. So foreground is kind of about like making

decisions and then the background is more about uh execution. And so just expanding on this, here's again uh more of a discrete mental model of like which task should go where. So for some of these foreground things exploring an API any sort of like spike UI tweaks that you're watching, those probably want like a human attention in real time. But for background tasks, maybe you're doing a refactor or migrating a a whole branch or a database. These are things where you can give the agent and this isn't to

just say that you spin up an agent and give it a one-sentence prompt and let it uh let it refactor your entire codebase. This is something where you spend a lot of time up front um figuring out the structure of how you want to do that. And maybe you're using like MCPs to pull in tickets from Jira or Linear and give the agent really specific context. So kind of the rule of thumb is if the output is reviewable or verifiable via an artifact, you can push it async. But if it's a bit fuzzier, that's when you want

to stay in the foreground. And then what does this mean for like an engineer's day-to-day? >> [snorts] >> So what we're seeing across folks that are utilizing these sort of cloud or background agents, um agents are writing pretty much all the code while the humans are spending more time like reviewing the artifacts like we talked about or coming up with how to solve specific problems. And then you can use agents to run in parallel to to kind of speed up the uh the capacity of what you can do in you know a day or

something. And so what this looks like is that you're writing a lot less code but being a lot more specific and uh you know maybe the job is turning a bit more into figuring out how you can give the agent all the correct context that it needs in order to run async.

Uh and then you're going to be reviewing these artifacts as opposed to uh actual code. And then you just want to ensure that you're doing this safely. So if you know you have multiple agents, you want to set up uh a like an environment and a branch setup that that you can do this uh in in a safe and in good way.

And then kind of like one of the last pieces here is just this sort of reshapes how you might think about meetings where instead of having like stand-ups or reviewing code in a meeting now people can show up to a meeting with a prototype and maybe you're demoing or iterating live with your team and getting really quick uh sort of responses and feedback loops from like the design and product people in your pod. Uh and so this is sort of another thing that's changing as the the

workflows of engineering are changing as well. And then we talked about this a little bit too but um tickets are kind of becoming like agent briefs. And so I think there was a lot uh of discussion when AI like first came out about prompt engineering and how you need to be good at prompt engineering and sort of the effect on that of the output that you're getting from these LLMs. And I think like tickets now are becoming ways that you can give really detailed information to the the agent.

And so here are just kind of some some best practices for uh how to ensure that the agent has everything it needs to uh carry that out successfully. And then this is pretty important uh for rolling something like this out, you probably don't want to just start with every single person on your team. Maybe pick a repo a few tasks a few engineers, start with that and then after you see some measurable impact, move to a couple other teams, uh share that up and scale uh within your organization. And then

once you see really good ROI on specific tasks, that's when you would want to sort of scale this out to the greater org. Um and then we can just share this after but just kind of some like do's and don'ts for using these background agents. >> [snorts] >> And then kind of getting into like how Cursor fits into all this. So um I think a lot of times people maybe still think about Cursor as uh primarily sitting in like this foreground agent world where you're writing shotgun and

you're watching the code being written. Uh and Cursor still is the best place to do all of that. So for foreground agents, you're maybe using that directly in the desktop app. You are giving feedback in real time to the agent. And then for those background agents, that's where you're going to use some of those other surfaces we talked about. Maybe you're kicking it off from your phone, maybe you're kicking it off from a Slack channel or you're using automations that

are integrated through any of the other like partners that we integrate with. So like a PagerDuty, Datadog um any of those other like CICD type tools. And then the last piece is we are giving you all the tools you need from like a governance perspective to actually control what goes on.

And then what this might look like. So this is just kind of like a diagram of how you might use Cursor to orchestrate a uh pretty complex like agent flow here where because Cursor is model agnostic you're using the best high-reasoning frontier model. Maybe it's GPT-54 to come up with a plan. Then uh you give you know a prompt here. So we're asking it to add 429 handling to APIs. And then maybe that main agent is breaking up that work into sort of three distinct

sub-agents that will run in parallel. And then once they're finished, they'll provide those artifacts back to the main agent. And then you'll kind of have that parent agent review all that work and decide whether or not you want to move forward or uh if there are fixes or things like that that need to be made.

So I think the visual here is just showing how you can use different models that excel at different tasks to optimize not only costs but also speed and ensure uh that you kind of have these things set up within your Cursor deployment to run automatically and it doesn't require a whole bunch of human oversight. So that is kind of it for the presentation today. Um we have a a blog post that kind of walks through sort of like Michael, our CEO's thoughts on all this uh that uh you can

go to by visiting the Cursor blog. And then we'll also share these slides. So if you guys do want to dive into anything uh that we sort of talked about in more depth, uh you can take a look at that. And then yeah, I would also just recommend taking a look at cursor.com/blog where we do have a lot of our engineering team write uh interesting pieces on things they're working on.

Think a lot of questions that we weren't able to get to today but if you do want answers to those questions, feel free to post them in like the Cursor forum and the like and and we'll make sure we get to that. Um and definitely encourage you to check out our automations page in our Cursor plugins marketplace on our public website. That'll give you a ton of insight on how the Cursor team is actually using cloud agents and automations today for a lot of our work.

Cool. Yeah, thanks everybody for coming. We'll see you on the next one.