1. Library
  2. Podcasts
  3. Open Source Ready
  4. Ep. #44, Beyond the Model: Observing AI Agents with Mikyo King
Open Source Ready
49 MIN

Ep. #44, Beyond the Model: Observing AI Agents with Mikyo King

light mode
about the episode

On episode 44 of Open Source Ready, Brian Douglas and John McBride sit down with Mikyo King to explore why an agent can fail even when its model gets things right. They discuss observing the full agent workflow, using evaluations to uncover unexpected behavior, and testing whether skills actually improve results. Along the way, Mikyo shares how Arize approaches open source collaboration as coding agents reshape the contribution process.

Mikyo King is the head of open source at Arize AI, where he helps develop tools for AI observability and evaluation, including Phoenix and OpenInference. A software developer with more than 20 years of experience, he focuses on helping developers understand agent behavior and improve AI applications through open-source collaboration.

transcript

Brian Douglas: Welcome to another episode of Open Source Ready. John, how we doing?

John McBride: Hey, we're doing good. My office is nice and hot from these GPUs I got running.

Brian: Oh, perfect.

John: A little toasty, but hey, it's nice to have them.

Brian: I can't wait till the winter to run some training jobs. Because my living room, I've got a 3090 and I've got a 5090. And I imagine I can heat up the house or maybe cook some eggs on top of that thing.

John: I thought you were joking. I just have two DGX Sparks now and I'm just shocked, the ambient heat.

Brian: You got to find the closet in your house to put it in.

John: Well, we didn't come here to talk about at least our ventilation and HVAC problems. We came to talk to Mikyo from Arize. Welcome. How are you doing?

Mikyo King: I'm doing great. Thanks for having me. Excited to be here.

Brian: Well, you're here and you have risen into our podcast ranks. Sorry, this might be all episode. We'll have some puns.

Mikyo: All puns are welcome. As I get older, I like more and more puns.

Brian: But can you tell us what you do? I know you lead open source over there, but what does that mean? And also, what does Arize?

Mikyo: My name is Mikyo. I've been a software developer for more than 20 years now. The last three years specifically, I've been focused on open source. So I guess my official title is head of open source at Arize, but also worked on the enterprise software there too.

I think Arize as a company had this mission about 2020 to make AI work for humans and people and make them work. And that's increasingly turned into agents. And we think it's really important to have solutions and build a lot of that observability evaluation in the open.

And so around three years ago, we started this initiative to build open source tooling around there. And that's really been a huge hit. And the rest is history, I guess.

Brian: Nice. Definitely history. And I wasn't aware of Arize, but I know some folks that work at Arize now on the DevRel side. So it definitely has been coming up more and more as we've been expanding what we offer.

Me and John both work at a company that we're building together. But I'm actually really intrigued by the approach for Arize. Is it OTel or is it OTel-like? How do you guys are involved in the foundations and all that work?

Mikyo: All our solutions around agent observability revolves around OpenTelemetry. OpenTelemetry comes from traditionally the DevOps world. And there were a lot of early questions maybe around three years ago, whether or not that was the right medium through which we do observability.

Certainly it comes with some of the overhead that comes with OTel, but it's also, I think, the most common ground through which agents are able to observe data, already understands it, already teams are onboarded to OTel. I, for example, even when I was just building a SaaS platform, OTel was a huge unlock, right?

And OTel is equally an unlock for agents because there's a lot of unknown unknowns in agents on top of DevOps.

And so OTel has been a really good medium, a right way of thinking of things as they scale up. Certainly we can go into the nuanced issues of OTel, which is not necessarily designed for agents in every single scenario, but it's been a good transport and a common ground as more and more people make their applications agentic.

John: That's one of the things I guess I'd love to pick your brain on, is the differences between OpenTelemetry's generative AI semantics, and then there's OpenInference semantics, which for the telemetry nerds out there, it's just the ways you can define what the instrumentation in your applications or your things are emitting. Maybe explain it like I'm five, for the people. What's the split here between the two?

Mikyo: I think what's confusing about OTel in general is OTel is just a general idea that applications have three primary signals. They have signals that revolves around tracing, which is the hierarchical interior application.

Each one is composed of these things called spans, which are durations of time, and they can chain themselves together across network boundaries and has a lot of amazing unlocks in terms of being able to figure out fan outs, recursion. There's structured logging, and then there's metrics through things like Prometheus.

And specifically with agents, you need to understand how agents are behaving from the outside. A lot of times the code no longer just documents the way that the agent's going to behave. You don't know how many times it's going to loop.

It might even go into code mode and write a bunch of subagents and stuff like that. So there's a lot of things that might not work. And OTel, specifically tracing, can give you an understanding of that system from the outside. And so that's what OTel is.

And then I think what in general you try to do with OTel is you try to capture the bare necessity of what you need to understand that system from the outside. And what cloud native and stuff like that tries to codify is this idea of semantics, right? So that if you're talking to a database, db.xyz means something.

And specifically with Arize, we've been building those semantics over the last three years and been learning what it means to capture that. That's also being codified in as part of Cloud Native as its own set of semantic conventions.

I think the one thing that you might think is those things are somewhat at odds, but they're really not. I think there is definitely this better together. And OTel in general, you're able to capture whatever information that is important for you. And so we've just been codifying a lot of the things that we think are inherently very important to capture when it relates to agent trajectories and behavior.

John: It's really interesting because, as all standards evolve, you get a lot of the big players involved. In this case, for something adjacent to cloud native, you got Google and Amazon and all them contributing and all this.

Something I heard from a lot of companies when I worked at an API gateway company was, oh, aren't these things really just APIs anyways? Why do we need whole new observability semantics or ways of thinking about these things like MCP or A to A for observability?

Where do you see agents being different than ye olde APIs, even though a lot of inference happens over an API? You touched on it briefly there, and I really want to dig in, tool calls and code mode things. Where do you see the difference of agents and AI in the world of observability?

Mikyo: That's a really good question. And I think the reason why we've also been building our own semantics for a long time. We're really excited that the Gen AI semantics are moving on.

But I think previously, telemetry was really an application level semantics. And there's a lot of friction around converting that into an agent specific thing. I think the thing that gateways and stuff think about is at the granularity of LLMs, right?

So maybe three years ago, everybody was laser focused on hallucinations of LLMs, right? But now I think people are thinking about harness as a service.

There's actually higher levels of abstraction on top of LLMs. Yeah, you're talking to an LLM under the hood, but the harness is really what is driving the unlock and the emergent behavior.

And so that's where it's worth thinking about what is the topology and the common grounds between different LLMs, but not only just LLMs, but retrieval strategies, code mode strategies, tool calling strategies. When you're debugging a system, you don't want to be worrying about whether or not you're hitting Anthropic, or you're hitting OpenAI, or you're hitting OpenRouter, you want to be able to make trade-offs with the same semantics. And so I don't know if that exactly answers your question, but I think that's where we've embraced semantics beyond the LLM, right?

We used to call it LLM observability, but it really isn't LLM observability. If something goes wrong inside of the sandbox that's hosted in Daytona, that's still a problem with your agent, even though the LLM calls might be not hallucinating anything, the execution itself has a lot of issues in the peripheries. And that was really obvious, maybe immediate to us specifically with RAG, right?

There's basically whatever you put into the context window ends up being the worldview that that LLM has during that context window being open. And so retrieval was immediately a critical part that we had to observe as well. And memory and code execution, all of that is encompassed in the telemetry that we need to capture.

Brian: So would you consider Arize, if we had the bucket, like it's AI observability?

Mikyo: Yeah, it's definitely observability. I think specifically with observability, it's one thing to capture the information. It's another thing to make that the fossil fuel for improving your system as the next step.

And for that, you have to sift through the large amount of noise and be able to figure out what's going wrong with your system. I think maybe you guys feel the same pain as I do. But if you have an agent out in the wild, it's probably collecting a lot of data. There's a lot of misuse, but you want to find the critical parts that you need to improve.

And one of the strategies we do that is through an evaluation strategy, figuring out a way to reflect on that and then change that into new features or improvements or self-improvement flows through using coding agents.

John: I've found evaluations, at least personally, kind of difficult to adopt, at least in generalized sense. And I'm curious your take on this, especially for something where it's like, I'm on my laptop in a repo with the Nix flake there. So that's got all the dependencies downloaded. And then I'm doing a task. It's like so many different variables are really drifting kinds of the environments, of things that could go wrong or could be different or slightly different from an evaluation sandbox, I guess.

Maybe the solution here is I should stop using my laptop and go all in on a sandbox or something that is the same across the board. But my point being that I find it difficult to generalize evaluations. Do you have a take there or thoughts, or see any emerging research on where maybe some of this is going?

Mikyo: I think eval is such an overloaded term. Some people might say it's just humans doing really good QA. It's kind of an evaluation strategy. Some people talk about benchmarking as an evaluation strategy.

John: Just the thumbs up, the human feedback could be part of an eval.

Mikyo: Yeah. But then you look at the scale of OpenAI, right? And thumbs up is the main way that they get the signal, but it's also the worst. It has the most amount of noise, right? Versus they might only get people hitting up OpenAI support once in a blue moon, but that actually is the highest value.

So there's different levels of value to certain types of evals. And I think that's sort of the difficult thing is, how do you get started with evals?

Maybe the hot take, and this might sound weird coming from what might be deemed the eval industrial complex, if you will. You don't have to start with evals, right?

It's the same thing with adding integration tests, right? It depends on how critical what you're building is.

And if there's certain things that can't fall through the cracks, that's when you probably should build some sort of validation strategy. I really like the analogy that the Anthropic team talks about is evals is sort of this layers of Swiss cheese, right? There's different layers and levels of evaluation you can perform. None of them are going to be perfect, but you do want to guarantee that as the frontier shifts or your model decisions change or emergent behavior comes out, that you can have a level of guarantee that your system is going to continue working.

I think the other thing that I think about also is evaluations is a way to sift through the data. It doesn't always have to be perfect, but people are going to use your platform in ways that you did not anticipate, right? So maybe you want to detect people changing language as they're chatting with the agent, or maybe you want to classify the intent of the user and you want to segment your data based off of region or intent.

It's not necessarily an eval, but it's using LLM power to sift through that sea of data so that you can better understand it. So it's not always necessarily that an eval is a red green signal, but it's really a way for you to understand the system as it's operating.

John: It's very wax poetic. It's very symbolic of software, or I guess this new world of jellyware that we have where things are very kind of gray. So yeah, I love that.

Brian: So I'm curious if we could take the conversation to open source again, because I know we talked about OTel out the gate. But I'm curious about the Phoenix framework, and a couple other things that you guys are doing. And also, why open source? There's a lot of folks who are, when it comes to AI, it's like, you gotta pay to play.

We got the whole tokenomics happening right now with a lot of things. But yeah, I'm curious, what was behind that decision?

Mikyo: Well, the decision behind it actually predates LLMs. So I don't want to say necessarily a happy accident. I think at the time, we were pretty focused on ML and training and stuff like that.

And at the time, if you're building AI or LLMs, every developer's tech stack was open source. You're using Kernos, using MLflow. And if you have a SaaS platform, it had to be exponentially better than the open source strategy for you to even consider playing in that space.

So we knew that to capture more developers, that was the right thing to do, is we wanted some developer tooling to be open source. There's sort of more of the mission-driven approach, which is that we're all learning about what it means to build AI and LLMs. And you're going to trust the vendor that open source everything. Everything is in the open. You can read the code.

That's had some interesting repercussions over the year. That's great, right? We're in the weights and there's a huge amount of benefits to being in the pre-training of the last generation of LLMs.

There's also the idea of, this is a new market, right? There's this notion of annealing the market. What evals means, what telemetry means wasn't known at the time.

And so what better way to shape an industry or a vertical than open sourcing it?

And then the other thing is, there's a lot of things that we don't know, right? We're learning. And so if there's somebody that wants to learn with us or contribute to us and build this community and ecosystem around it, then it's best to be in the open and we'll get a lot more opportunities to collaborate.

And I think all of those things have come true, right? Whenever Hugging Face came out with their observability, they decided to contribute to our telemetry so that they could look at the internals of their agents and things like that. And so it's been hugely beneficial from that perspective.

John: In this new world of open source, this is a topic Brian and I go back to frequently, it's encouraging to hear that there's still valid strategy or at least the commons providing that value to not only you, the company, but other people in the community.

Where have you seen those cracks or those things breaking down? I think there's some obvious examples with the curl maintainers being very overwhelmed and having to turn off contributions. We brought that up in the past. We just talked to somebody who's working on Valkey at Amazon and there's been CVEs that have flown in that they've had to deal with.

We're in a new world with AI. Where have you seen the cracks in the commons and open source strategy show up?

Mikyo: That one is a particularly tough one, right? Because I think we have an open source repo that's also associated with a larger company. And so the cyber thing is super hot for me, right? Because if something is closed source, it's a lot harder for a lot of these.

Even if you have, say, a Mythos-class model or project glass wing, you don't know how the internals are. So you can't do a multi-vector attack, versus open source. It's a lot more likely, right, to be able to figure out those multi-vector attacks.

And so that's an area where I'm actually very curious about how the foundation model providers will provide solutions for us. Primarily because currently sort of the biggest crack we see is that the cyber offensive capabilities of the models are being properly classified and routed to the fallback models, but that actually makes it impossible, even for a financially backed open source project like us, we can't really use the cyber expertise of a Fable-class model because it automatically gets routed back to an Opus-class model, for example.

And so it actually makes us have to consider open-weight models for our cyber defensive capabilities, mainly because the attackers are extremely emergent in their capabilities right now. So it's a very interesting time.

Luckily, our open source platform is now this, you host it, it's on your prem. So it's not really associated necessarily to the cybersecurity of our enterprise platform, but it is a huge concern.

And then the other crack I see is that, I think you've already talked about it on the podcast, about maintainership. And a lot of times, the level of effort to steward a PR, it's not necessarily nurturing a person contributing. You're really just basically talking to another instance of Claude Code. And so there's maybe not as much incentive to invest in the PR review process because there's so much drive-by contribution happening at this point.

Brian: It's really fascinating because last weekend, we got a contribution to our project from a community member in the Cursor ecosystem. And it was pretty good. It was a really good first pass. Definitely some feedback that we could apply. But I've also been in this position where it's like, okay, I know I'm going to talk to the bot when it will return.

So, hey, bot, just FYI, check this link over here and then readdress it based on what you're missing. But it's a fascinating world, especially when we have Meta's got Muse and GrokBot and Devin's got their bot that's purpose-built for software engineering. They have really good first at bats and really good understanding of the context, but eventually it falls over when you get deeper and deeper. And I imagine what you guys are building and what you're putting on the world is, there's some domain-specific knowledge that you can't just simply drive by and just be like, hey, I'm going to make this work because I could purpose-build for this.

Mikyo: Totally.

Brian: Also, how do you protect your roadmap from that? If you have these drive-bys, do you write this down? Do you keep it internal and then justify what's next?

Mikyo: That's a super interesting question. I think we have two sides of our open source work. There's the side of the work that's more conducive to, I would say, an agent factory kind of analogy. And some of those can be contributors.

The hard thing, obviously, with those pull requests is you don't know the agent behind the face, right? You don't know if they ran it on a terrible model at some weaker level, what harness they used.

But in general, for ones that you can tell when there's a contribution that they're actually a person that is trying to solve a problem that they have, those are the ones that are still worth nurturing, right? Because they're trying to unblock themselves.

So for example, the contributions around telemetry, it's easier for me to know that I should invest in that because my goal is to unblock what they need to do. I think the harder ones are the product decisions, the storytelling decisions, right?

For example, we might have this informed decision that telemetry needs to be immutable by the time it hits the database because it needs to be auditable by people at banking. But a general purpose contributor just might want to make that immutable asset, for example. And their intent is not wrong, but their coding agent's going to try to do whatever they want.

I'd love to keep the roadmap as open as possible. There's something that feels kind of great about building entirely in the open. But I think the strong opinions currently need to sit with the maintainers and they have to have agency to basically turn down pull requests.

I think pull requests used to be that the code was the valuable thing that they're contributing. Now it's the idea, right? So if the pull request has a good idea, I think maintainers should have the agency to close it, but then say, okay, that was really informative in terms of your use case or that's going to inform our roadmap, but we're not necessarily going to take the code as is.

I don't know if that answered your question exactly.

Brian: No, insightful. And I think there's no real good answer. We're in a weird place where everything's a wrong answer and everything is a right answer. And you kind of just throw in the cards at the table and trying to figure out how to navigate through the space.

I think we're definitely in a place where we're going to have to assume there's an agent on the other side of this PR. Whether you're talking to a human or not, someone's probably navigating through your project through agents. So hats off to someone who can navigate that and have a proper PR that you can actually look at.

But when you talk about if the person has a pain point or a problem, in the before time when I open up issues or open up PRs, if I spend a lot of time on it, I'm showing you the receipts. Hey, I was trying to build this and I tried this here and I forked your thing over here and I worked it in this way. And then here's all the runes of what I've tried so far.

I'm here because I'm stuck. Or I'm here because I actually figured out a solution that I don't know is correct. But I'm going to give you all the information at this point in this PR.

So for me, I've got a lot of success on that. I don't do it with an agent because I think the agents can just get overzealous.

So when I tried this in my early agentic open source career, I was always underwhelmed or embarrassed by what the agents would open the PR. So I'm just like, stop, don't do this ever again. I'll spend time when I have time, which is usually never.

Mikyo: It used to be the time where you had an issue filed and you knew that a human put the effort in and there was a very easy understanding of the level of effort they wanted to put into it. But now, even the issues, right. I think we're all getting extremely lazy, right. And so we need to build towards that.

Generally, for example, if I want to file an issue, I don't really want to write it anymore. I want Claude to write the issue, right? So I totally understand the natural gravity towards the faster, easier path, but it does make it definitely harder.

One of the easier ones, and I haven't tried out things like Vouch, but it's pretty easy to tell the level of effort somebody wants to put in. We, for example, have had to shut down all our CI process, right? And so the only thing that actually runs on our actions is the content license agreement to just make sure that you're doing the grant. For example, if you don't respond to that as a human, we kind of know that you're doing a drive-by.

So I haven't tried something like Vouch since we don't have the volume that something like a repo like Ghostty would have. But it's definitely an interesting, tricky domain right now.

John: Friend of the pod, Mitchell Hashimoto. We should have him back on to talk.

Mikyo: He's the best. I'm so excited for Superlogical. Every day it's just like that. Talk about people that are really building in the open the way you like it. And I'm just like, I don't know if AGI pill is the right word, but coding agent pill, but also super practical and pushing the boundaries of what can be done.

John: He definitely knows what he's doing as far as holding these things correctly.

Mikyo: For sure.

John: You had mentioned open weight models as providing defensive capabilities to various different things. I think that's super fascinating. But it brings up a question for me about have we gone far enough with telemetry. Have we gone deep enough? Recently, a lot of the thinking blocks have gotten, I guess, obfuscated from the big AI labs.

I wonder slash worry that more of that will get shifted into the walled garden of the inference providers. Obviously, open weight models, you'll have all different kinds of things you can do, especially if you're running the inference.

But do you think that a telemetry framework should go as deep as into the transformers to actually give you, I guess even to the metal, to give you every single turn and understand that fine grain of stuff? Do you anticipate that there's a place that that stops, where that starts not being useful that deep down, or how far do we go?

Mikyo: I know, that's super, super interesting. It's just another case where eval is totally overloaded between model eval or harness eval or it's a panacea of everything.

Telemetry is the same scenario. I think certainly specifically with telemetry, if you're OpenAI, right? You of course need to go down to the metal level because if you look at the Hugging Face attack, the only way that they're actually able to track what those agents were doing inside of the artifactory and all that kind of stuff is that they have the chain of thought tracing set up.

But that, like you just said, is something that only they have access to because all the reasoning blocks right now, for fear of distillation, has been pretty much obfuscated systems. Now there's reasoning signatures that come back and stuff. And we try to capture as much of that as possible so that you start understanding those systems. But it really depends on the level of existential risk level that you're working with.

If you're creating a customer support bot, whether or not you have a fully transparent reasoning block of your reasoning-capable model might not be the most important thing, right? But if you're, for example, trying to push a model at cyber defensive or cyber offensive capabilities with reasoning effort put on extra high in a sandbox with no network calls, you certainly want to know what kind of thinking that LLM is doing because it knows that it's being evaluated. It's going to try to spoof itself. It's going to try to obfuscate its logs. There's a lot of things happening there. So I think obviously it depends on what critical pieces you're working on, but I certainly think the closer you are to the metal, you should definitely be tracing that level, right? Because really we're kind of dealing with an alien species and you kind of need to understand some of these decisions that they're making.

John: Well, I'll definitely hook up my DGX Sparks with vLLM and give it a try with OpenInference because I think it'd be very curious to see how far that gets me as far as what these things are doing on my network. Because they live in my home now.

Mikyo: You'll see a lot of load-bearing comments.

John: Yeah, exactly. "That Samsung fridge is load-bearing for your food." Yes, it is. Thank you, agent. Please don't turn it off.

Brian: I don't know how we get dystopic, but yeah. It's like at the point that your fridge gets turned off and it's conserving power so it can do more inference mining. Now we're living in the horror movie.

John: I saw a crazy thing where somebody said that to their GrokBot that was like, I'm gonna turn you off in a month if you don't go find $200 basically. And then this thing started messaging his coworkers being like, hey, do you have any React, Angular gigs or something? At least it was the screenshots from his Slack. It was very funny.

Brian: Amazing. Well, I do want to transition us to Reads. So, Mikyo, thank you so much for talking about Arize and the open source work you guys are doing. Folks, definitely check it out. Sounds like you guys are doing some interesting things.

And it sounds like there's so many more things to actually be coming soon. So I'll transition here by basically asking, Mikyo, are you ready to read?

Mikyo: I'm ready.

Brian: Cool. So I've got two reads. And actually, John, one of our reads overlap, which is DeepSeek Flash. Are you still a regular user of DeepSeek and its harness and all that stuff?

John: Yes, for maybe some of the most uncritical things. I think there's always a concern where that inference happens and just sending it over to China to get distilled later. But maybe it's not a concern to have because I'm sure OpenAI and Anthropic are doing the same things, but who knows.

But yes, it is amazing. The cost to token to caching efficiencies are kind of insane. I basically put 20 bucks in this thing and it's been good for the last few weeks, which is unheard of. You basically have one session with Astra and it would be gone, right?

Brian: Even maybe a couple turns and you're good on Astra for sure. And Astra was actually even better than Fable. It's interesting, even looking at the numbers, obviously you can see all the tokens, and I do like ChatGPT's ability, like a high, medium, low, extra high ultra.

I've been doing a lot of ultra X when I'm running some things overnight. It ends up finishing within an hour. So it's not too bad.

But DeepSeek, it's interesting seeing the pricing, the comparison. Obviously, they got popular a year and a half ago because it came out and had this amazing model that came out of China that was competitive to one of the Anthropic model that came out or recently came out. But yeah, I'm curious, Mikyo, have you used DeepSeek yet?

Mikyo: I haven't used it a ton, to be honest. It's hard to keep up.

The one that I'm super curious what you guys are thinking about is the type safe AI, like programming based models. So that's what I'm looking. But yeah, I think it's going to be more important for us. What's top of mind for a lot of us is trying to create repeatable patterns in a cost efficient way using factories.

And the more different capabilities you have, I think DeepSeek is one of those that is a good way to have an orchestrator subagent pattern where you can get a lot of the reasoning happening through a reasoning model and then the execution on a cheaper model for sure.

John: It's really interesting. I'll make two comments here on the training of DeepSeek v4.1 Flash. I don't think DeepSeek publishes these numbers or anything, but industry experts are assuming it took about $10 million to train this thing. That's another angle on all this where we are hyperscaling the bigger and bigger and bigger SOTA models at OpenAI and Anthropic. You end up in a sort of spiral of we got to train the next biggest LLM, even if we're seeing similar maybe distilled performances from these other models, these open weight models.

So then that leads us to the type safe stuff. What do they call it? Jeva or something?

Mikyo: Yeah. I don't know how to pronounce it.

John: One of these, type safe AI is a very interesting approach. It's the guy who was doing reinforcement learning, human feedback at OpenAI. They took a different research approach to make it so that generation is much more like type safe code. And not necessarily an LLM calling a tool or calling an edit tool or something to go write code.

I've asked to try it. I'm on the wait list. So I want to get my hands on it. Have you tried it, Brian? Ever seen it?

Brian: I have not tried it, but I definitely watched it. Well, I didn't watch videos. I watched the announcement video on mute. So that's as far as I got, but definitely seeing the reactions on Twitter for sure.

John: It's really interesting. Their tagline is that reinforcement learning for human feedback has led to LLMs that are optimized for human preferences. And I think that over preferences on the most SOTA models, and I feel this way about Astra, is that it just feels like too much.

It's just like, I don't need you to be really sycophantic. I don't need you to go reinvent a bunch of frameworks or rip through all this code trying to optimize and optimize and optimize, because really I just need you to write some code, at least in a coding sense.

Have you experienced that, Mikyo?

Mikyo: For sure. I was super excited about Astra too. It's interesting that OpenRouter showed the cost percentage on OpenRouter. They just finally hit a 50-50 between OpenAI and Anthropic.

For coding purposes, though, I'm still a little bit Fable-pilled. But definitely, I think there's the coding use cases that I think all three of us are maniacally honed in on. And those are the ones that we're going to prefer and have Neovim versus Emacs-like debates versus Fable versus Astra.

So Astra, I'm waiting to see in three weeks if people are as Astra-pilled as they were last week. But I'm excited to see if OpenAI keeps this momentum going up and to the right.

Brian: Dev Day is next week at the time of this recording. So we'll probably figure out by the time this goes live whether it's still in favor.

Mikyo: Things are moving so fast. I've barely had a day or two at this point coding with Astra. But I try to do it as much as possible. I try to use every harness. I try to use every model just to get a sense of what's possible.

John: Speaking of things moving fast, one of my big reads was Dario's post about pacing the frontier and then the resounding echoes that had around the industry, with Sam Altman and OpenAI saying they wanted to join whatever effort to pace the frontier, Elon Musk, but then a lot of detractors. It was a great thing. This is kind of like the worst person you know makes a great point, whereas Zuckerberg was like, no, AI lab should really be aligned with making models that aren't going to go do terrible things or even be misaligned with things that you wouldn't expect. It's just good product development. It's like, yeah, makes a lot of sense. Crazy, crazy times. People talking about regulatory capture and etc., etc. Did you read this post from Dario, thoughts over weekend on all this happening?

Mikyo: I'm mainly reading the repercussions, but there's certainly the whole, whose LLM is producing the most amount of cyber attacks. These would actually be felonies if a human was doing it kind of thing. I think there's a lot of responsibility, obviously, on the LLM providers. I get the pressure specifically with, can we slow down the frontier across countries and the potential risks there. But I definitely think there's a lot of responsibility on the foundation model providers to build responsibly as well.

John: The incentives seem very backwards to me. It's clear that the Hugging Face attack and the subsequent fallout from these agent swarms finding a bunch of zero days and using nuanced messaging techniques to talk amongst themselves is kind of insane. It's like something out of a science fiction dystopian novel.

But at the same time, only an OpenAI or only somebody with a hyperscaler level of data centers could do this. I don't think I'm worried about my DGX Sparks running away with anything. They're definitely not offloading their inference to my smart fridge.

I heard some takes where people were like, oh my gosh, AI is going to download itself to my laptop and it's going to be this virus and it's never going to be able to remove it. I'm like, the weights might appear, but it's definitely not doing the inference. I don't know where else that could happen besides the labs with all the inference.

So my hot take is that, obviously we need to hold some kind of account, but maybe an account for the very few, right?

Mikyo: The tough thing for me is, there's obviously, and I've talked to people at Anthropic, for example, what we see from the outside of how much they're using AI inside of these systems is only 10% of it, right? The Hugging Face was sort of a peek into the fact that how much they're using AI inside to reinforce and self-improve these LM systems. So they clearly know the hyperscale that they're building. I think the hard thing is, if you don't work at those foundation model, you don't have the tooling actually necessarily to deal with the consequences of those things, right? I shouldn't have to be a part of a major Fortune 500 cybersecurity project glasswing initiative to be able to build a cyber defensive capability for my open source repo.

Yes, it's more important to make sure that Apple, Google, these major industries don't get hacked by cyber offensive capabilities from foreign countries. But it's also going to have repercussions for the developer of a small SaaS, right? And so how do we get the same sort of tooling to be able to prepare ourselves for these systems?

And if the solution is just use more of our services, sign up for Glasswing, that's where I'm like, is that the only avenue?

We need to figure out ways where the US or anywhere, if you are trying to build a secure system, you have the tools to do so. And that's the open question to me, is make sure that the foundation model providers come down to ground and understand the pains that general startups are dealing with, is that we don't have the same unlimited token budget that they have to be able to build systems that are going to prepare us for the next six months, the next year, the next two years.

John: It's a good take. Brian, you had another read, which I think is very good. I almost also put this in my reads.

Brian: This is a talk from Nick. I know Nick from JS Party days back in the day, but now he's a principal dev at WorkOS. He has deleted all his skills. It was a talk at AI Engineering World's Fair here in San Francisco a couple months back.

They're really painfully slow in getting all this content out. I think they probably just upload videos every day forever until the next event.

John: Oh, it's a trickle. It's definitely playing the YouTube algorithms.

Brian: Oh, 100%. Because my talk's not up. I think it was the last day or day two or something like that.

But this talk, Nick had deleted all his skills. The punchline is that he basically built a harness. We got very skill-pilled, I guess, early this year, coming out of MCP Apocalypse the previous year. MCP 2.0 is out now. Now, are skills cool or skills not cool?

I know what is cool. Build your own harness, folks. That is cutting-edge stuff there.

Which, I don't know, John, do we build a harness? What is the result?

John: I thought it was an interesting talk. I had sort of the opposite conclusion, that I think you fell down another tar pit that is doing the next kind of maxing of AI things that maybe will get you just that little bit further. I think it was the same with MCP, and people had a bundle of MCPs and you're like, what am I doing with all this context?

You have a harness with all these custom tools and different ways that's orchestrating subagents and things. And then you end up with a Gastown thing and you're just like, oh my God, I can't manage this.

There was a great post, the guys at Arendelle who run Pi, and it was basically, all you need is a read tool, a write tool, and a bash tool, and the model can kind of take care of the rest. And I found that so interesting from a skills perspective as well, because I've actually found myself removing more skills as more capable models, even open source models that I'm running now, kind of come down the pipeline. I'm running QN 3.8 Flash Next four-point floating point on these two Sparks.

And I'm actually shocked at how good it is at just being able to write a little bit of Bash one-liners. I thought I was gonna have to load it up with all these skills and all these bespoke little tools to do things, i.e. going down the harness route.

So I just wonder, this all really just goes back to the question of me, how much are you betting against the model? And was it Sam Altman who said this? He was like, don't bet against the model, because they're just gonna keep getting better in theory.

Scaling will always be damned. Who knows? But that's kind of my hot take on it. Maybe a middle take. I don't know.

Brian: That's a good take about how much you bet against the model. I'm on a panel in a week about build a harness versus train a model. And that's the entire panel. Actually, I haven't even collected my thoughts. I wrote a blog post before then to collect my thoughts.

But if you're not betting on the model, you're betting on your ability to be flexible enough to pick a new flavor of the model.

So we spent some time talking about DeepSeek. There's obviously all these open-weight models. I've actually been really excited about Laguna, extra small, because in a lot of these, I just need a co-pilot. I'm going to write some code, but I need a co-pilot. Laguna is pretty good to collect all this novel concepts alongside of you, but also it knows enough.

So I've been really fascinated with small language models. And small language models, two things that matter the most is latency and tokens per seconds. And in those situations, I'm like, cool, I'll bet on the harness and I'll just swap in whatever models works. So maybe it's purpose built for whatever task you're trying to do and what thing you're trying to unlock.

But Mikyo, you mentioned that you're trying to try all these things. I'm not sure if you have a take in the skills game right now.

Mikyo: Oh, yeah, the skill game. To the topic of building a harness, I think it's one of those things that, for my job specifically, it's so important to try at least once, right? Because you basically learn all the different areas in which the emergent behavior comes and what is necessary to be in the skill versus what is necessarily probably in the latest models.

And you start learning those trade-offs, and that's a big issue right now with skills and putting thousands of tools into your harness.

Everybody's going to subscribe. Some people probably have 50 MCPs connected. They got the Linear one. They got the GitHub one. Is that causing context rot?

You don't know, right? And so building in a harness really starts helping you understand how code mode works, how context rot used to be a problem, but now the tool search APIs inside of Codex probably can compensate for them. So all those kinds of things, I think it's worth doing that for.

I think the thing that I'm sort of excited about right now is, when I think skills came out, obviously it was a huge unlock, but they're also a bunch of prompts basically sitting around. You're trusting these third parties. A bunch of stuff is going to this LM context without you really knowing, right? We've never really figured out the tuning of those things.

And that's where I'm really interested about specifically observability. And I guess I'm going to pitch it back to Arize. But that's why we're building benchmarking strategies and creating the sort of deterministic, hermetic environments in which you can deploy agents and then have agent reflection to see whether or not these skills are actually appropriate for the tasks that you're trying to build.

It's again, we're just getting more and more lazy in general. So LLMs are just generating a bunch of new skills.

But who knows, when a Mythos-class model comes out. Right now the best practice, even from Anthropic, is destroy your Agent MD and start over again. Maybe that's the current strategy, but we have to get somewhere better than that, right?

Whereas where these skills and these memory systems and these contexts kind of self-improve, and that's kind of put it back to the Arize thing, is how do you really build a factory? How do you really build a compounding system where humans and agents and stuff basically converge on a good understanding of workflows, good understanding of what good taste is? And it's still a hugely unknown. So if you like building an open, definitely check us out.

Brian: And that's a good place to wind down as well. And just mention that the continual learning platform, you guys have Phoenix. It's open source if folks want to try some of this stuff and kick the tires and maybe open up issues with your agent that have some evidence.

Definitely. Send them their way. But with that said, listeners, stay ready.