
Ep. #12, Shipping at the Speed of Understanding with Steph Hippo
- Platform Engineering
- AI
- Developer Experience
- Continuous Delivery
- Observability
- Site Reliability Engineer (SRE)
On episode 12 of Third Loop, the Progressive Delivery team speaks with Steph Hippo, Platform Engineering Director at Honeycomb, about what happens when AI accelerates development beyond the pace humans can absorb. They explore safe defaults, evolving code reviews, and the feedback loops that help teams understand what they ship. Steph makes the case for reinvesting saved time in better decisions, software quality, and human connection.
Steph Hippo is the Platform Engineering Director at Honeycomb, where she leads the storage, data ingest, and site reliability engineering teams. Her expertise spans platform engineering, reliable production systems, developer experience, and helping teams adopt AI while preserving technical understanding. She previously worked in site reliability engineering at Google.
transcript
Adam Zimman: All right, we can let the fun start.
Heidi Waterhouse: Woo-hoo!
Adam: Before we get too far into the fun, something that we want to make sure that we do is we have a wonderful guest with us here today. Steph, if you'd like to introduce yourself, you can tell us who you are and where you're currently entertaining yourself and any other fun details that you'd like to share.
Steph Hippo: Yeah, I'm Steph Hippo. I'm the Platform Engineering Director at Honeycomb. I'm based out of Seattle, and I just really love making software systems run better for everyone.
I am a big soccer fan. I've got two young kids and a dog that will be at my feet for this podcast today. But I'm just really happy to be here and chat with y'all.
Adam: That's awesome. Well, thank you so much for joining us. We really appreciate it.
You were one of the folks that came highly recommended for this podcast from Charity Majors. She said that you would be a great person to talk to about all things progressive delivery. And as we've been talking about this third loop, where we really want to think about the aspect of user adoption from the perspective of software development.
And something that we think is increasingly important, and I'm sure that you and your role at Honeycomb have seen this play out time and time again, where it's not about what you build, it's about whether or not people use it. Because it turns out you can build all sorts of cool engineering stuff, but if your users don't use it or aren't actually getting value out of it, it's probably wasted effort.
Why don't you tell us a little bit about what your role is at Honeycomb and how you spend your days?
Steph: So as the Platform Engineering Director, I am responsible for all of the teams that help run Honeycomb at its core. So our storage team, our data ingest team, SRE. And so we are doing all of the things behind the scenes that make Honeycomb keep humming. We want to make sure that as we're scaling and as the AI landscape is changing, we're able to keep up.
And so a big part of the heart of reliability is, I'm sure you've heard Charity say, nines don't matter if users aren't happy. And users can't be happy if nothing is there.
It's really about can we continue to be there for customers at often their toughest moments. An outage is often why people are coming to Honeycomb. And so we want to make sure that we're resilient and that we're able to respond in timely fashion and be able to do so now in cooperation with your robots.
So it's a whole new class of users, and it's the wild west right now, but it's been a lot of fun figuring out what works best, and it's changing every day. And I'm sure by the time I sign off here, there will be a new problem to go tackle.
Heidi: So before we get too far into it, describe for us what you think platform engineering is, because you sort of touched on that, but it's controversial. Is it building tools for developers? Is it the stack? What do you think of it as?
Steph: Yeah, sure. Everyone puts their own definitions and stuff into the platform word box. Computer science has never had a problem with overloading terms ever.
But for me, it's really about making it easy to do the right thing by default for developers.
And in this case, it is integrated into, like I said, sort of the backbone of Honeycomb that every product team needs to be able to build successfully and integrate with the whole picture. We want to make it really simple for our teams to build and scale in ways that it's very difficult to fail in novel ways.
We always say we want to learn things each time, right? And so we want to fail in new and exciting ways, not in ways of lessons we should have learned already. So that's the whole ecosystem that we look at end to end to try to make that harder to do.
And our users are both Honeycomb's end users, but then also our internal developers that are in turn serving those end users. So it's my favorite part of engineering. I often joke the best and worst part about it is that all your users can message you directly internally. And so it does a lot with keeping a tight feedback loop and making sure that you're responding to their needs.
Adam: Quick question for you. Maybe not a quick question. But how has this changed as your ICP has split recently, right? In the sense that now all of a sudden you have human users and, as you said, the robots.
Has this changed the way that you think about either the prioritization of what you're working on or the way in which you architect and build a feature? Do you actually think about it in terms of this we are building for robots versus this we are building for human eyeballs and fingers that will need to manipulate things? Do you make that active segmentation?
Steph: Everyone's favorite answer: It depends. Right? I think at the heart of it, we're building for humans.
Honeycomb wants to make it so that software engineers have the tools that they need to understand what's happening in production, whether that's an outage or whether that's trying to understand how to incrementally improve your systems. And so more and more, engineers are bringing robot buddies along with them. And so we do need to be able to build for their robot buddies.
And so there have definitely been patterns where we now have to think about what the agent interface is going to be like. And there are a lot of things that we can take advantage of there too. In terms of Honeycomb, being able to query your telemetry data, you had to understand your telemetry schema pretty in depth to be able to form the best, most performing queries. Now we can teach an agent to do that and you can do more natural language forming of the questions that can be then answered by Enigom.
And that's been really awesome, both for developers and I think leadership too. I think we're seeing more and more folks outside of the engineering function get to understand their production systems as well to ask these questions.
And it's really lowered the barrier to entry to folks that might not be technical in that way previously that they can now ask and get their own answers.
And so designing for that has been a ton of fun, and it's been really cool to watch it unlock so much information for other roles.
Adam: In that way, do you try, or do you actively, are you successful in looking at the user interaction of the features and functionality that you're supporting and rolling out and being able to actually discern between human users and bot users? Is that something that you look at? Is that something that you can tell the difference on? Is that something that you track actively in terms of how you think about future development?
Steph: Yeah, we have a ton of folks that think very deeply about our design and user experience at Honeycomb. And we have noticed they use the tools slightly differently.
Ultimately, I want to build software that's going to augment human teams to make them better. If we do get to a point where the robots have it all now and they can take the pager and we don't need to worry about it, that's great. All the humans can go bowling is my thing.
But we're not there yet. So for the meantime, we're absolutely developing for both.
Adam: No, I mean, thinking of it more in terms of when you're thinking about your building new features for the platform for your developers to be able to build on top of. I know that increasingly the developers at Honeycomb are using agents to be able to do agentic development and actually working through building code, building things like that. So I'm guessing that there's aspects of the platform where you want to make it more ergonomic for a robot to do the query or to do the request than it used to be maybe when a human was doing it.
Steph: Yeah, so more on the internal users for sure. I do heavy dogfooding at Honeycomb. It's everywhere Honeycomb uses Honeycomb. Probably shocking no one. So we form a lot of those opinions: hey, we wish it did this.
Or we recently asked everybody, what's a feedback loop we don't have yet that you wish we had? And then, do we need to go and build that?
Heidi: That's a great question.
Adam: What was a good answer?
Septh: I think there have been a couple good answers. Things looking at, okay, I want it to both help me instrument the change or feature that I'm pushing out, and then also update its own harness, right, to look after I roll it out to know, or do we have to pass, fail, did I just break prod, right? And we've had other feedback loops of, hey, we know that this area of the code base has been maybe neglected over the years.
We want more feedback loops that can go and look at, hey, do we have dead code there? Hey, do we need to add more tests there? Is our test coverage burning us? Or even things where, because the agents can do more natural language processing than before, we're looking now, hey, what organizational feedback loops haven't we closed in past incidents and postmortems?
And so trying to look at not just the technical pass-fail, but also what do we need to learn as an organization and as a team of humans. What's not having the intended effect right now? And so that's been really interesting, especially as you read more about the risk of the cognitive surrender or pushing the limits of just the cognitive load of the humans. How do you use AI to put some of that back in, right? And to be able to exercise your knowledge of the system that's now changing faster than ever. And we know it's not the same as just getting a digest of what changed at the end of the day, right? It's not the same thing as reading and interacting with it and exercising your brain to understand how the systems are built together. And so we're also considering those feedback loops of adding those learning opportunities back in. We're very heavy users of Dr. Cat Hicks as a learning opportunity skill and that we have heavily adopted internally at Honeycomb.
Adam: We're big fans as well.
Steph: I'm sure.
Adam: One of the things we have also been noticing, and we've talked about previously on the podcast, is the way that teams are starting to spend more time up front on the kind of architectural conversations, not what can we do in terms of capacity or writing new code, but starting to think about it in the context of what should we do. We can do a lot of things because we have this additional capacity from an agentic perspective, but we're also need to be now even more discerning with, is this actually good for the user?
Thinking about it in the context of human change acceptance rate, there's a certain threshold that you're going to hit where, if you change too much too quickly, your users, to Charity's quote, will not be happy. So how do you think about this for your internal users? I joked in a blog post not too long ago that developers, on the one hand, will tell you, oh, the system should be organic, it should be changing daily, and things like that. And then I pointed out the fact that developers as a class of humans up until recently, the vast majority of them were using an editor that was built in 1976. So it's like, change the things that I want to change, but don't change the things that I don't want you to change.
Heidi: Well, change the things that I don't use, but don't touch my tools, which I think is super interesting because we never, not never, but we have trouble thinking of what we're building as somebody else's "our tools."
Steph: I do think for a group of people that hates having their cheese move, they're really excited to move a lot of teams. What a time to be alive, right? That we are now at the point of software development where we're like, oh, well, maybe now we have too much change too fast, right? I remember when I worked on software teams that were trying to go from shipping quarterly to shipping weekly and then daily and getting the release train going. And now it's like, we can make it do whatever we want almost whenever we want.
And now it's like, okay, maybe we can actually start looking at other bottlenecks. And it's okay if the end user's ability to absorb the change is the bottleneck. I think we're going to hit human cognitive bottlenecks. And for an industry that loves going faster and faster, it's probably going to have to learn a bit more discernment now.
In terms of developing for developers, I think that we try to give a lot of choice within reason. There is complexity that you add if there are too many different ways to do things, right? And so like I said, we want to make sure that you can kind of fall into the pit of success.
If you're trying something out for the first time, the failure mode should still land you somewhere safe that is doing the right thing that will help you keep moving. But we know that folks like the different tools, they have different ways of thinking. And so we'll probably offer a few different options of ways to do things depending on your role and what you work on.
I think one of the conversations that we've really focused on over the last year was how each team within Honeycomb has to be able to fit with their risk curve in terms of AI. And so my storage team's probably not going to turn the agents loose for free for all, right? They're very responsible that way. I love that about them.
There are other teams that have been able to experiment quite heavily because the cost of a rollback or something like that is much smaller. And it's been really cool to see what they've been able to iterate on much more quickly.
And so if you're a team that is working in architecture that has a lot fewer two-way doors, then you're going to be more cautious, and that's okay. That's not a failure or refusal of you to adopt AI or anything like that. It's having good engineering discernment.
Heidi: It's understanding your pace layer.
Steph: Exactly. Thank you. That was the phrase I was looking for.
And so it was making it a challenge on, which group do we serve first, right? And so when we're talking about what can we deliver to the end user, it's a little bit easier to work with the stuff with the smaller blast radius, you can iterate faster. And then we have to work our way up to letting the agents get into those deeper pace layers. But we're seeing it and we're adding those feedback loops and finding ways to make progress there. But again, it's the very complex systems. And if the limit that we hit is the cognitive load on the humans, cool. Again, what a time to be alive.
Kim Harrison: So you're talking about how you have some customers where people outside of engineering are now using the product. So there's going to be people who don't want everything changing all the time, but there's going to be new change that's incredibly exciting. Tell me more about that side of it, some of what you're actually seeing.
Steph: I think this is where natural language as an interface goes a long way. I'm somebody that did not love natural language interfaces at first. There are still things I'm not thrilled about. Again, give me my editor from the 70s.
Heidi: What if I could get a straight answer?
Steph: Right, exactly. But looking at how our sales teams, our marketing teams are using it, or some of our user research teams, just being able to ask in natural language, hey, what's going on here? Hey, this particular customer, how many users do they have? Is this feature making them get more value out of Honeycomb, less?
We also have other signals that are starting to show up. We're like, okay, folks are querying a lot more, reaching for it a lot more often, but also their sessions are getting shorter because they're getting the answers that they need and they're leading. Cool.
Adam: That's a win, right?
Steph: Yeah. It's not up into the right cost of U1 to be measuring for delivering value. That's just been really cool to see.
And so when we're designing for these other users, I think the problems that I'm running into now is the all-you-can-eat buffet of AI is kind of over and usage-based billing is coming. And so making it easy for them to get the right model for the right task, right? I don't want them to feel like, hey, I can't touch AI anymore because it's getting too expensive. I know this is again where I need to design for you now to make it easy to reach for the right thing at the right time.
And wherever I can, I actually want to start building tools where you don't have to make that decision. We can just do it for you.
But it's also some enablement and teaching them a little bit about, hey, here's how this stuff works under the hood. And this is why this is more expensive than this. And that also helps them talk to customers who are having similar problems. And so I think it's broken down some of the walls.
Adam: On that point, how are you thinking about that when you're designing the systems? Are you doing something where you're actually doing kind of pre-processing or triage to a query and saying, hmm, this looks like a small, medium, large or easy, medium, hard problem, and then determining which model you're going to actually send this request to, or how do you manage that internally?
Steph: Yeah, so we are doing some query segmentation and classification that helps us make sure that we're using the right resources to get you the answer that you need. But it's also for folks that are doing things other than asking queries. It's just kind of our internal operational usage of AI.
Are you using the most frontier expensive model for your tiny little questions, or are you using it for your planning stage and then condensing down into maybe a cheaper lower level model?
And we have limits now on what folks can spend, but we're pretty lax about giving you more. We just want you to use that limit as a check-in and be like, hey, tell us some of your usage patterns. Can we help educate you a little bit? And then can we learn about what we need to build for you so that you don't hit these limits?
But it's not a stick. It's a carrot to come talk to us, not, how dare you use so much of this, because we want people to still get as much value out of it as they can.
And so we're thinking about those kind of feedback loops in the message that we're sending to users. Hey, we're not saying no, we're just saying...
Adam: We're just saying stop using Fable to tell you what to get for dinner.
Heidi: Yeah, or who was it? Somebody was like, a huge portion of this is going to translating PDFs to slides. Please, no. But can you please explain how these four variables are interacting to cause this weird thing on the flame chart?
Steph: Exactly. And so being able to help guide people through that, we did so much that trying to build their reflex of just when to reach for AI. And now it's like, okay, now when to reach for which AI. And so I think it's a natural part of the evolution of it. My job was easier when we didn't have to worry about that, so.
Adam: Fair. I mean, harder in different ways, right?
Steph: Yeah, right.
Adam: So with your teams, when you do planning and stuff like that, do you find that you're spending more time or less time doing planning versus then sending the team off to do execution or building? Have you noticed a change or a difference there?
Steph: We have been writing a lot about this on the Honeycomb blog, actually, about how as we're using AI more, you should use the time that you win back to talk to humans more.
And so I think it depends on what you put under the umbrella of planning.
Adam: So what do you put under planning?
Steph: Yeah, let me tell you what we put under planning. So a lot of it is the reflection and how engineers feel about how sustainable their job is right now, right? We were definitely getting feedback that folks were feeling fried by the end of the day because they were now keeping so many more threads open with AI, not always necessarily closing all of them. And then having some team time reflection, like, hey, is that actually serving us or not? And so kind of planning how the team plans, right?
So I have one team that's been great at running different social experiments on the team, like, okay, I think Charity mentioned some teams like, we're not going to use AI for two hours as a group and see what we need to talk about. Or we're going to do other things that let us explore a bit more on this particular area, and we can spend more time planning and thinking through what we might want or what we think is feasible.
And so it's definitely different than figuring out, okay, we need to find what we can launch and iterate and ship and get out the door, right? And so now it's getting us some intention back, I think, and letting people breathe a bit and making those more one-way door decisions that they have to live with long term. So there's been other things that have been great and have made planning much easier. It's been awesome to write up a ticket, especially for things like a little paper cut, and just hand it off to a bot. And also we're doing much less planning that way when we can just say, bot, find time to do this in your day. And it's helping get through some of the smaller polished stuff that we otherwise would have to wait for when somebody had a free minute.
Adam: And how are you dealing with code reviews for those types of things?
Steph: Oh, what a great question. There have been a lot of discussions around code reviews, the purposes that they serve, and how much control we want to give up to the robots.
I do think now there are robots that are better at code review than even some top engineers.
So if, again, you acknowledge that code review always had purposes other than just catching bugs, right? It was being able to absorb those changes to your system that you're accountable for, being able to understand other people's work and how everything kind of tied together. And so that's where, okay, if we are going to take code review and put that more on the bots, where do we still get that system level knowledge share? And it's probably just shifting left, right? To the planning stage. And so now are we just doing more design review and prompt planning than anything else?
Heidi: Are we writing better specs?
Steph: That's the dream, Heidi, right?
Heidi: I'm like, wait, this sounds strangely familiar, where we plan everything exquisitely and then it just runs on rails.
Steph: Yeah, exactly.
Adam: I think that this is something that I was actually getting at with regards to spending more time in planning. And this is the kind of planning that I was wondering about.
Actually, coming back to that question or statement I made earlier, the planning sessions used to be, can we do this? And thinking about it in terms of, who has capacity on the team? How many days or weeks is this going to take? Where is this fitting in in a sprint or things like that?
And now what I've seen is a shift to, should we do this? Or at least a lot more emphasis on the should, right? And thinking about it in the context of the outcome that you're actually trying to create for the user. And focus on the problem that you're trying to solve. It reminds me of Heidi jokes of getting a better specification and then just having things execute. There was a lot of that that was supposed to happen, even with waterfall, right? Or with agile, where it was this was the whole premise of, no, you spend the time up front to build the plan and then you can execute more smoothly.
And I think one of the challenges that happened with those types of software development life cycles was simply that the execution timeline actually was too long. And what happened was that the needs of the outcome actually changed, or new information was brought in that wasn't known. And it was like, okay, well, there needed to be a better way to adjust.
And this is where developers in particular saw that need for a change or need for an update to the specification as they started building. And that's where the continuous delivery started to get in traction. It was like, okay, now we're at a point where all of a sudden, we have the ability to execute on the code development faster than the external needs or observations or changing the specification.
Steph: I also think part of why those systems started to no longer meet the needs of engineering teams is systems also got more complex. And as you got more complex and you had more complex interactions, you were going to need to have finer specifications. You're going to have more feedback loops.
And so I feel like we've been building towards this for decades, right? And so now the future is here again, which is time to be alive. But there are still things that we want to do.
And now I think the bottleneck is going to be your ability to make decisions as an organization and stick to them long enough to execute on them.
And so in a way, that's always been true. But I think the speed at which we can execute now is going to really highlight an organization's ability to get everybody on the same page and get everybody's concerns addressed.
Adam: So one of the other guests that we've also been fortunate enough to have on the show is Chad Fowler talking about the Phoenix architecture. And that was another recommendation from Charity. But I think that this brings up some of the things that you're talking about and encountering with regards to cognitive load.
And how do you start to create, whether it's actually a microservices-based architecture, or at least this notion of pace layering that gives you the ability to segment out and have clearly defined boundaries of services or objects within your stack that allows you to focus on what is the input that's coming through and what is the output that we want to create for this thing? Are you seeing that as something that's coming up in conversation for you all of looking to actually reduce the scope of a particular service or aspect of your stack?
Steph: Yeah, I think, again, because so much of the focus has been how do we deliver more faster, then you've seen the complexity build up.
And so we've seen engineers lamenting that AI has added a lot of complexity to their systems. And I'm like, cool, how do we make AI put it back? Put the simplicity back.
And so there have definitely been conversations on what do we need to do now. I think that goes back to the planning and being more intentional. And so what do we want to be true?
One of my staff engineers says, we can make Honeycomb do whatever we want. And I was like, this is true. What do you want to be true?
And there are parts of the code in every company where they're like the haunted graveyards. And I'm not going in there, right? It's like, well, now we can send the bots in. Or maybe make it less daunting.
And so we are certainly thinking about what do we need to simplify. What are the interfaces that we can make better? What are the things that we know would make our lives better, but have been putting off because it's either too tedious or doesn't deliver immediate value to an end user, things like that.
And we've had a couple migrations that we've finally done where we've taken it the whole way to the end. And instead of letting that last 20% of the migration take 80% of the time, you can fully push it across the finish line. And it's been wonderful. And so, yes, I would say there's things like that.
And then re-architecting comes with its own risk and its own pace layers. But there are lots of ways now that, again, we can add a feedback loop: what would you want to be able to feel safe about this change? And we can spend more time budgeting for that work that's going to increase your confidence in delivery of the change.
And so I think it will take a while to get people to decouple AI with speed and now be able to focus on AI with quality.
Adam: In that context, another guest we had was Kent Beck. And he was talking about one of his more recent books, Tidy First, we touched on. You're shaking your head like you're familiar with this concept.
Steph: Yeah.
Adam: Is that something that you actively push the team to do? It sounds like a little bit of the same ideas where you're actually saying, hey, look, before you start touching any existing code or doing any new feature that touches an existing service, do you actively look and say, are there things that I should clean up in this code or this service first before I look to expand it?
Steph: Yeah, I'm a huge proponent of this. I liken it to moving houses and apartments years after college, moving every year. My husband and I have been in this house for a little over a year now. But we found something where we're like, this has been expired in the pantry and we've moved it.
Heidi: Twice?
Steph: Yeah, at least. And we're like, that's embarrassing. Why are we doing that? And there's the software equivalent of this all over the place.
I actually came to Honeycomb to help build out its, we were calling it Apne Woman at the time, but they were responsible for the design system that had been a tragedy at the commons and we'd had a lot of front-end tech debt. And so we were like, okay, we were a design system team without a designer yet. We were hiring for one. So I was like, all right, so the designer gets here. We are going to clean up as much as we can.
And it was phenomenal just how much you learned about your problem set there, cleaning out the cobwebs and saying, oh, I didn't know this was back here, kind of thing. And then we were able to drive a lot of simplicity and make the system just easier to understand just by cleaning up. And then you can make better decisions when you can see everything.
And so we now have a design system that we are quite proud of. I don't run that team anymore. But it's incredible that we're at the point where it's like, hey, should we even still be investing in this?
Could we actually put this down for a little bit and go somewhere else? And what a time to be alive. But it's been phenomenal to see more of that work get prioritized.
Adam: Bring in Marie Kondo every few weeks and just say, "if it doesn't bring you joy."
Steph: My dream software developer, honestly. Get her in here.
Heidi: I think one of the things that I'm most interested in for platform is how much you want the platform to be obvious and how much you want it to be invisible. Do you want people to think about using the platform or do you want it to just be background?
Steph: I think you need to be able to abstract the right things at the right time. So going back to, hey, we moved from waterfall to agile to whatever you want to call this now. The reason we were able to do a lot of those jumps was abstraction and simplification, where eventually developers didn't have to think about a layer of problems they had to before. It was kind of a solved problem in the platform.
But until you have solved all of it, you need to expose some controls on that to your developers. When I was working on platforms at Google in SRE, one of the things we said were like, hey, we could automate some of this, but then there are some decisions that are made here that we can make guesses and reasonable defaults. But I think we want the developers to be aware of the decisions they're making. And so maybe we don't pull those yet until we know that they are now boring enough choices that it is okay.
And so that is where I think there's more of an art to it. And that's where a lot of the platform user design comes in.
You have to earn the right to make it invisible.
And so giving those choices, and then once you feel like you are at a point where either it's sufficiently solved a problem or you know that there are other ways for your developers to internalize what some of those defaults are, then you can do it.
And it's not always clear and obvious, but I always think about my operating systems professor in college. He was like, yeah, it used to take us 30 days to artisanally install an operating system. And that's how we liked it.
And I was like, cool. I don't really need to touch my operating system day to day anymore, right? And so we'll see the same thing with different agent harnesses and things over time.
So, yeah, I would love to make more of it invisible and take more cognitive load off of my teams. But you have to incrementally work there and earn the right.
When I first started in SRE, what they drilled into us was know your service, know your service, know your service. And as things advanced, it was like, okay, know your platform, know your platform. And so once you knew your platform, you could drop into different services or applications.
And even if you didn't work in that team day to day, you would have enough guideposts from the platform to get your bearings what the service did and help take care of it, right? And maybe it's just enough to mitigate, restore the service, and then pass it back off to the dev team that has the deep knowledge of the business logic to fix whatever went kaboom and try again. And so that was how they were scaling SRE.
And so I think you're seeing the same principle now where it's like, okay. But add a new abstraction layer. Exactly. And so that's exciting.
I can't wait to see what developers don't have to care about in five years.
I have a four-year-old right now. And I'm like, is she going to have to learn how to drive? I don't know.
Adam: In fairness, just so you know, I was asking that question when my kids were four. They're now 17 and 20.
Steph: And driving?
Adam: No, neither one of them are. That's the funny part.
Heidi: Have to learn to drive and get to learn to drive are two different problems.
Adam: Two different things, yes. But they're just like, no, public transit works great. And they live in a city.
Steph: I mean, that's my dream. I would love for them to never drive. That's also why I live in a city.
Adam: But I hear you. I do have a little bit of concern that the learn to drive aspect is a little bit like jetpacks. It's forever 10 years away. That's kind of where I'm at right now.
But I agree with you. It's things like that: what are some more aspects that we as humans are either A, terrible at, or B, don't enjoy doing? And so how do we put the robots to task on doing that stuff?
Steph: And somebody will have to remain an expert in it for when somebody needs to pop the hood. And that's why we do that now. That's why we have specialists, and I remind engineers all the time, you are not expected to know everything at every layer because no one does.
Even the developers that you most admire, they can't just jump between layers of the stack. I'm sure that they would have the resilience to maybe eventually learn and figure it out, but they're not going to know it offhand.
Heidi: Tanya Reilly says everyone's back end is someone else's front end. And it's so true. It's like the things that are obvious to me, and then the things that are too easy for me to care about, and the things that are too hard for me to deal with. And it doesn't matter where you sit.
Steph: And that's why we have teams, and it's why we have organizations and specializations. And so that's how we all tackle it together. But yeah, Tanya is 100% right.
Heidi: I was talking to a sysop friend of mine, and she's very clear. She's a sysop, right? She's not SRE. She's not platform. She runs servers, man.
And she said, automation seems great, but it's always fragile. Because anytime anything changes, your automation becomes wonky. So I was wondering how we're going to deal with that going forward if platform sort of abstracts a bunch of that away.
Steph: Great question.
Heidi: Do you have a theory?
Steph: Not yet. I appreciate your insight. I mean, I think that is where our continued investments and how people learn and understand their systems are important too, right?
And so I've been talking to my teams again, going back to the code review question. So if code review has all of these other things that it has given us to understand our systems and socialize understanding between humans, where else can we put that back? So things like, hey, is it easier to run a mock incident now for you to learn from? Is it easier to get people onboarded and ramped up by walking you through different exercises, things like that?
Probably. Are we making time to do that as a team? No? Okay, let's change that, right? But yeah, I think that will be the million dollar question: how do we evolve our other team practices to be able to keep up with that?
Adam: No, I mean, I think this is where it definitely is clear. We're still in the messy middle. And to your point, there's so much that we're still figuring out. And the way that we're going to get there is, it doesn't matter until it's actually running in production.
We're not going to learn anything, just keep it in the lab. Or the learnings that we get are not going to be nearly as impactful. And I think that that's always been the case.
Steph: Yeah. And just making sure that you're still valuing learning, I think, is all you can ask for right now. And then I hope it gives people a sense of job security, right? There's always just a new class of problems under the next abstraction layer.
Adam: Yeah, absolutely. And I do like the idea of taking the reclaimed time and spending it more with human communication.
Heidi: Is there a difference for you between learning and curiosity?
Steph: For sure. Charity talks a lot about this. And so when AI was first taking off and people were still very skeptical, myself included, we were asking folks to try it out because if you didn't understand the tools, you weren't going to be able to build for the end users that were using them. You weren't going to understand where they needed the safety gates and things like that.
So even if folks had a lot of very reasonable objections to how AI came to be, now we're like, okay, the box is here. How do you want to make it safe? And so you have to be able to engage with it.
You're still allowed to be a hater, for sure. You do need to be an informed hater. And so you maybe won't double your productivity with AI, but can you double your curiosity for me?
And can you go see, what is this hype about? I will say I do think there was a tipping point last year. So I had a baby last year before I went on maternity leave. I was like, AI is still causing way more problems for me than it's solving. What is everybody so excited about?
And then had my baby, didn't do any AI with him. Came back from leave. And I was like, wait, did this get good now?
Adam: No, AI didn't change. It was just you having a baby. That's what changed. Haha.
Heidi: Yeah, it leveled you up.
Steph: Robin was like, hey, now you're going to get it. But yeah, so it was me trying to continue to update my curiosity about it.
But I think what was tough for a lot of engineers is because of the rate of change and how quickly things were getting better, right? We had always talked about what makes an MVP. And so if somebody would put a product out there too early that didn't meet the M yet, you would leave a bad taste in your mouth as a user. And then the bar for going back and trying again would get higher each time the product disappointed you, until you give up and you go to a competitor.
And I feel like there's a lot of that with AI at first. And engineers were like, why would I try this? And I just tried it last week.
And it let me down. And so I'm like, last week, that was a year ago. You need to go try again.
And so finally, I do feel like we are leveling off on the rate of change there. Maybe now that AI has heard me say that, it'll take off again. But it really has done a lot to break some of the human expectations of, okay, you have to try again.
You have to re-engage. You have to stay curious because it's like, oh, what's the AI up to now, kind of thing. Is it up to no good or is it about to supercharge me? I don't know yet.
And so, yes, I would say curiosity and learning is going to reign supreme. And again, it always has, even more so now. But I don't know.
I'm enjoying that I told you so there. It's like, just try things and learn. It's fun.
Adam: Awesome.
Kim: Who else should we be talking to?
Steph: I can hand you loads of people. Would really recommend Katherine Cass. She has been working in the agent space since before it was cool and focused a lot on continuous delivery and release pipelines. And so I think she would probably just nerd out with you very hard for an hour.
There are some SREs that I think you should talk to. Jennifer Petoff. She's wonderful.
I owe so much my credit. She is so good at thinking about how do you teach and train people in terms of helping them understand complex systems. And she's a huge influence into why I became such a good SRE.
And then Fred Hebert from Honeycomb does a lot of our systems control stuff. He's a principal engineer, one of my SREs. He's very into the socio-technical systems of software delivery and understanding complex systems. And if you give me another five minutes, I'd probably generate a long list.
Heidi: Those are some great names. And we really super appreciate you taking the time to talk with us and unpack some of what happens when you're thinking about platform as part of how we touch users.
Steph: Awesome. Well, thank you. I had a lot of fun with you. And yeah, I'll go let my dog inside now.
Content from the Library
Is the Key to Understanding Code Treating it as a Graph?
Why Code-as-Text Doesn’t Work for Genuine Understanding Major engineering orgs claim that as much as 30% to 75% of their new...
Why On-Device Inference Needs Custom Observability
The Unique Challenges of Mobile Compute A significant focus in modern AI has been on large language models with billions of...
How to Make Agents Durable for Concurrent Systems
How Can Agents Be the Future if They’re not Reliable? Some reports suggest that as many as 25% of enterprises have adopted...




