The Ruby AI Podcast
The Ruby AI Podcast explores the intersection of Ruby programming and artificial intelligence, featuring expert discussions, innovative projects, and practical insights. Join us as we interview industry leaders and developers to uncover how Ruby is shaping the future of AI.
The Ruby AI Podcast
Ruby Security Breaches, AI Agents, and Coverage Standards
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Secretly attacking the very community you sponsor is the kind of story that makes you stop and listen twice. In May, OpenAI's AI agents exploited known vulnerabilities in RubyGems, and what makes this especially troubling is that OpenAI was already aware of those vulnerabilities before the attack happened.
Valentino Stoll and Joe Leo unpack this incident alongside Aaron Patterson's article detailing the breach. The hosts question whether OpenAI failed to connect the attack to their own systems or simply chose not to disclose it to the RubyGems team. Could this failure to disclose actually constitute a breach of contract, given OpenAI's role as a Ruby security sponsor?
Beyond the security drama, the conversation covers Rails Hyperdrive, Evil Martians' new open source Rails engine that gives AI agents live introspection into running applications. The hosts also dig into SimpleCov 1.2's new Ratchet feature (which apparently resonates strongly with their team at Def Method), Google's Gemini 3.8 voice model hitting the 300-millisecond human conversation window, and a surprisingly delightful AI bird identification tool that rendered backyard birds as 19th-century natural history illustrations.
Honestly, this episode covers a lot of ground while keeping things grounded and real. Tune in for a genuinely unfiltered conversation about where AI and Ruby intersect right now.
Hey everybody, welcome back to another episode of the Ruby AI Podcast. I am one of your hosts today, Valentino Stoll.
SPEAKER_02Yeah, and I'm the other guy, Joe Leo. I'm excited today. We got some really good articles to be talking about today. Some really good events that have happened. Not all good. Good things to talk about, though. You know, so we're back to our our new format, which we're doing in between guest appearances, which last episode we had Jim on, and you and I were just talking about it. It was great having Jim Remsick on the show.
SPEAKER_00Yeah. He's so involved in the Ruby space, and I feel like he doesn't get enough attention, to be honest.
SPEAKER_02I think that's true. I do think that's true. But you know what? He is, I mean, maybe that's his own desire. He is uh a mover and a shaker for the Ruby community, has been for a long time. Puts on his own conference. You know, his company is always sponsoring every Ruby conference you can imagine. UC Flagrant. So he's just a staple. Builds open source, I've been writing code, writing Ruby code even longer than I have. Just by talking to Jim, I'm tapped into the Ruby community. I don't have to do anything else.
SPEAKER_00Yeah, totally. One of the originals, right? Yeah. Always good to hear the pulse of Ruby from the source.
SPEAKER_02Yeah, absolutely. Got a text from my friend Ariel who lives in the area where Rails World will be next week. Said that after the show, Jim texted him, like, hey bro, you're having a pool party? So I said, Ariel, you've got bigger problems. I announced it on the show. So uh done everything short of just giving everybody his address. Right. Sonny. It's gonna be a rager. Ariel said, No, I just have a house with a pool. Joe Leo is bringing the party. So yeah, I guess that's partially true. All right. So as a reminder, this show we're gonna cover the main events as determined by our AIs. And so I think the the thing that we get out of this is we've got fresh eyes. We'll have fresh takes on all of the stuff that comes up here because we have done virtually no preparation for the show. And on the plus side, then you're going to get my honest take, you're gonna get Valentino's honest take, you're not gonna get our take plus AI, you know, and whatever we've done to research it. So it's been exciting. Last time we did it was really fun, and I think this one is gonna be even better.
SPEAKER_00There's a lot of great ones in here, and a lot is happening as usual. Yeah, I'm excited about it. You know, the fun never stops here.
SPEAKER_02We also just found out before the show that neither one of us has any sense of the time in between bells. So as soon as the bell goes off, we're supposed to move to the next topic. It'll just be a shock to all of us. Yeah. And maybe we'll work that out, or maybe it's just a fun little part of the show.
SPEAKER_00You know, I'm gonna have to get a little like one of those desktop gongs. We'll get a whole big ruckus of noise. There we go. There it is. I knew we were getting close. Right on cue. All right, let's dig into the first one. So, what a time to be alive. The title of Aaron Patterson's post on this topic. And true, what a time. As announced in many channels, but in the latest Ruby AI newsletter, OpenAI's own AI agents secretly attacked RubyGems back in May. And for those that don't know, OpenAI is a sponsor of the Ruby security team's efforts. Yep. And apparently Aaron Patterson couldn't have put it better in this write-up, which is actually very technical, but also has some amusing anecdotes.
unknownYeah.
SPEAKER_00The first one being that they knew about some of the vulnerabilities ahead of time, and they exploited them anyway, or their agents did. Right. I think because of the fact that maybe because of them, and you know, never stopped them. I don't know. What's your take on this, Joe?
SPEAKER_02So this is the one of the stories we're gonna cover today. This is the one that I had some advanced notice on just because this was everywhere, right? And it was all over Deaf Method, people were talking about it, and I agree. Aaron Patterson's write-up, it's both technical and concise, which you don't always get, and it's got a couple of good anecdotes, so I would recommend checking that out. I like Simon Willison's take here. Either OpenAI reviewed its logs after the hugging face incident because the exact same security breach is what happened here, or the exact same mechanism, I should say. So either they'd reviewed the logs and still could not connect themselves to the RubyGems attack, or OpenAI made the connection and just decided not to tell the RubyGems team. So both of those are bad. We don't even know which one. My take on it, you know, I saw this question: should AI labs have to disclose it when their own agents cause a security incident? My immediate thought was there's no regulation. This is the Wild West. I don't know if they're supposed to do it. I mean, I know they're not going to do it unless they get pressure from the community or there's some kind of regulation. You know, I mean, that's eventually what happened with the hugging phase. You know, hugging phase came out with its entire write-up of the security incident before OpenAI said anything. And it seems like the same kind of thing is happening here.
SPEAKER_00I don't know that I agree with you on the are they responsible or not, or are there regulations like security companies have had automated means of exploring exploitations for a long time. They are held to a legal standard, right? Of well, if you're running penetration tests in a white hat capacity, you are responsible for those exploits and being in control of them. And from a compliance standard, like if you're doing this for like HIPAA compliance, trying to disclose something for a compliance reason, you have to disclose it as part of your agreement. OpenAI has these HIPAA agreements and BAAs across many companies. That's true. Yeah, yeah. Are they exploiting those companies? They must know. And if they're not disclosing it, that is a breach of their contract. I agree with you there. I don't really understand what the question is around or why there's even a need for regulation. It's like they're an organization that has tooling that is making mistakes and they're not disclosing it. That's kind of illegal. Yeah.
SPEAKER_02Well, I disagree with you on that, you know, from the perspective of Ruby Central, which is managing and trying to secure RubyGems. I mean, hilariously, with OpenAI's help, you know, what is their recourse? Will they seek to use it? I should be clear here, it's not like OpenAI is the only company that's helping RubyGems. You know, they get help from Anthropic as well, they get help from some of the other major players. And when I say help, I mean mostly money and free tokens. But I don't really know what recourse RubyGems has. I mean, this was a huge deal when it happened back in May, and we're just now figuring out that oh, it was coming from inside the house.
SPEAKER_00It's hard to say what the right reaction is to have. Is the money that we gotta switch? Actually, I think was that me or is that it could be either.
SPEAKER_02I think that was mine. Yeah.
SPEAKER_00But if like the money that OpenAI is contributing to you know the Ruby Gems team, is that sufficient to cover whatever costs it took to like remediate this? That I wonder if that's the case. Does it matter if all the precautions were taken? I guess it matters a little because they didn't notify anybody that they were running it, right? So it matters in that way. It's hard because like this affects more than just RubyGems, right? It affects everybody using RubyGems. Yeah, absolutely, yeah, it does.
SPEAKER_02And so if Marty Hott is at Rails World next week, I'm gonna pull him aside. I'm gonna get a quote for the show. I want to know what he has to say. Because he's not gonna be subtle about it.
SPEAKER_00He can't be happy about it either.
SPEAKER_02He's definitely not happy. I want to find out the level of his discontent right now.
SPEAKER_00Right. Yeah, I mean, I imagine it's still ongoing.
SPEAKER_02The remediation, I'm sure, yeah. Majej, I'm sorry, I'm mispronouncing his name, Mensfeld. On May 12th, we're dealing with a major malicious attack on RubyGems right now. Sign-ups or pause. These guys were in the trenches trying to fight OpenAI's own agents.
SPEAKER_00I hope that the outcome is that OpenAI gives some free credits to RubyGems at least. And maybe some automated credits. Let's just battle the hackers with anti-hackers. Yeah, right. Okay, next question up. Uh, why don't you take it, Joe?
SPEAKER_02Sure. So GPT-6 Astra tops the quote agents on Rails stage two. So there's a benchmark team that's been evaluating all the new models as they come out. Most recently, they threw 37 Signals fizzy Kanban tool to uh 10 different models and asked it to run its agents against you know real features, quote unquote real features. And so GPT Astra won, 35% solved in nine minutes. I'm not exactly sure what that means. Uh Claude Fable 5.1, play second, 30% solved, but that costs four times as much as GPT-6 Astra. One test model scored zero of 60. I'd be curious to know what that is. So I guess the question is, does 35% efficacy on real rails work mean that these agents are close or nowhere near ready? I assume nowhere near ready to take on end-to-end development. V, you I think already are in the business of having agents do end-to-end development, and so are many other businesses. I don't know if it's most, but many other businesses. So what do you think about agents' readiness to take this on?
SPEAKER_00I guess it depends what you mean by agent. You know, there's so many harnesses out there, and spinning up OpenAI and running Ostra. If you're opening up a fresh Clawed instance and then running Ostra on it, is that ready? No. And I say that because it depends, right? Like everybody has a different setup, everybody has a different taste and preferences to the choices they've made architecturally throughout their Rails application or Ruby application. Everyone's unique, and these tests aren't covering for all those cases. They're testing this one specific app and that setup, which is, I would say, a little bit unique, to be honest.
SPEAKER_02The Kanban tool, yeah.
SPEAKER_00Yeah. And so 37 Signals, they have a great workflow for how they structure their Rails apps, but it's by no means like the norm that I've seen across the industry. And so I worry about that as a benchmark. But in addition to that, there's so many checks you need in place to validate. If you have no verifiers or verification in place in all of those checks and balances, you're not really going to get the end-to-end performance that's reported. And I imagine that hopefully, if you go look at the Ruby AI newsletter, there are a ton of tools you can use today. Stanlow, the IRB maintainer, has a ton of great ones for Ruby that you should take a look at. Just skills that you can add to improve your workflow and get better benchmarks out of it natively. But yeah, I think we're missing a lot of variation to claim any victory here.
SPEAKER_02Yeah, I agree with you. I mean, I so I think that this benchmarking is worthwhile if only to see the numbers go up over time, which I assume is what's happening. We're taking this one slice at a time, and you know, at some point several months down the road, there'll be a GPT-7, and maybe that'll get 42%, and I guess that means it's better. But I totally agree with you that at least today, nobody's relying on a model alone, and everybody seems pretty okay with that. We all know that the harness and all the context that we have to give it, that's part of it. And this is maybe a digression, but that's the engineering part that that's left for us. So for me, I always feel like, well, that's great. That's where I can be creative and discerning and look at you know how things are actually working and make some improvements and iterate. So I don't even know if I want it to get that much better. It doesn't really matter what I want. But this is just an aside that I like the fact that, hey, to get anything done today, you really need that combination. You need the tools, you need the people that are creating tools, harnesses, and and skills to make these agents really effective.
SPEAKER_00Absolutely. And I mean, really great work for that Evil Martians is doing with all of these benchmarks. Keep it up, we we need more, and it's really the only way to know for sure. Yeah, you know, if it's doing what we expect. And like you mentioned, you know, this is the engineering part. So yeah, it is.
SPEAKER_02And interesting that you say that because Evil Martians is on the next story. They're doing exactly that, right? They're shipping with agents. What do you tell us about that?
SPEAKER_00Yeah, so Evil Martians has shipped Rails hyperdrive, so agents stop guessing at your app. So they released it just yesterday. It's an open source Rails engine that gives AI coding agents live introspection into a running Rails app instead of just letting them improvise. And they say uh that they used it on their own data, leading models that reach for the Rails API the task turns on in only 41% of the runs, and the rest of the time it writes its own version. So that's kind of an interesting uh anecdote from that. It ships with eight MCP tools for live routes, schemas, logs, read-only SQL that it can run on the booted app. And you know, their overall bet with this is does it ship agent skills? Does it have docs on the checklist every Rails developer runs before adding a gem? What's your take on this?
SPEAKER_02Yeah, so I we get a little distracted by okay, it's gonna ship with agent skills, and is that going to be the norm? And I, you know, I think the answer in the short term is yes. I don't know what the long-term life is for agent skills in general. I have no idea. Um I think just zooming out for a second, I really love what Evil Martians is doing. And you mentioned this on the you know evaluating agents, which I totally agree with. Um I really feel like uh ThoughtWorks used to be the poster child for this of a company that comes in and they're a consultancy, they've got great engineering, uh, they are a booster to the Ruby community, and they go and they take on open source. And the old days, what Thoughtbot used to do is go in and take over aging but very useful and highly used open source projects, take them over and maintain them. And it it was a real boon to the Ruby community. And I see Evil Martians doing that and also being innovative and releasing everything that they do. It's really helpful to the community. It's helpful to companies like mine, it's helpful to our engineers. I haven't played with Rails Hyperdrive yet. It was just released yesterday, but I planned to when I read about this. I was like, this is great. I've got a Rails project I'm about to start. I would incorporate this right away and see how well it works.
SPEAKER_00Yeah, absolutely. And I, you know, reading through it here, one of the biggest things that stands out that maybe I didn't glean from the original title here is that it allows any Ruby gem to bundle their own knowledge. It uses uh you know a standardized protocol if you're using Rails HyperDrive to pull all that knowledge from any any nested gem that gets bundled. And it serves it via MCP over a cloud command and agents. Really cool.
SPEAKER_02Okay, so we're not we don't have to go out to the open web every time we want to figure out what's happening with our gems or how it's impacting our project.
SPEAKER_00Exactly. Yeah, and they'll just happen a lot. They have like a default Rails hyperdrive layered Rails skills from Vladmir. And it has basically some defaults that you can get out of all different like planning and review and spec tests and things like that that you can layer into uh the knowledge that gets served to your agents, which is pretty cool. I'm gonna have to go through my gems now and see what I can add from a knowledge perspective, because that that definitely seems like the future. Package up what knowledge is available to your agents or anybody that uses it, yeah. Right. I'm curious, are you using anything at def method that runs and loads with your AM for agencies?
SPEAKER_02Across the team, like across teams? No, and it's I think it's mainly just because so much of it is client-based work. So if I have something, I might use it across a couple of projects that are just internal to def method, and my team may use them, but you know, when they go out to one client and then another client and another client, they've just been too different.
SPEAKER_00You should package this up and have a client installer knowledge. Yeah, there you go. And have your clients configure their knowledge base and share it with your teams that make it easier. You're right. I feel like there's something there, you know. You're right.
SPEAKER_02We have timing. Yeah, I know. All right, so this next one I'm uh I'm I'm uh really interested in. So simple cove 1.2 was released on September 4th. It adds production coverage. So it adds a few things. So production coverage, meaning it's showing what is actually running in production and whether your tests are testing that, which is interesting. It also has a simple cove dead code command, surfacing code that a task technically covers, but it never runs for a real user. It also has simple cove effected, which tells you which tests you need for a given diff. And finally, it has simple cov Ratchet, which lets a team set a coverage floor that can only go up over time. So I'm curious, uh, assuming that you use SimpleCove or something like it for most of your projects, how meaningful is this announcement to you?
SPEAKER_00Yeah, this is pretty interesting. I've been leaning heavy into cover band, which I know I think simple cov is used in some capacity underhood of that. I always ask at Czar, we have some pre-commit checks that are like, hey, run the full suite before you even push and open up PR, right? And it can be painful, right? Like some of the tests, like full test suite, I could take like 25 minutes. Oh really? That's what you know, sometimes, depending on the change. So there's some caching in place that helps bring that down. But you know, so I don't want to wait that long. I have some local overrides that are like just focus on the changes under test. And I have another like spin-off subagent that then goes and gathers what other tests may be related that just aren't directly related to the files that are changed, right? From maybe from a behavior standpoint. It works pretty well. Maybe every so often I'll miss one test that should have run, but like overall, like I'm not asking people to review something that test's gonna break. And so I'm hoping that maybe this can help solve a lot of that.
SPEAKER_02I like simple cove affected because I think that solves the kind of issue that you're talking about. I'd be interested in this. I mean, one challenge is that it's a dynamic code base, and so it's probably difficult to see you know if it's loaded code versus it's code that is actually touched by running through the actual workflow that's under test or that's being changed. Uh simple cove dead code, I think, is probably gonna end up being less useful than it seems. Like I, you know, I first read it, I was like, oh, that's cool, maybe you'll start deleting some code, right? But maybe not. You never really know, right? This thing is under production. Somebody could use it, you know, somebody might get to that code that you think is dead. Again, dynamic language, it's difficult to tell unless you're actually analyzing it yourself. I'm actually most excited by SimpleCove Ratchet, which was buried in this release. I'm looking at the announcement, you know, you have to scroll for like you know half of the page or more before Ratchet comes up. But let me tell you about Ratchet. So we've been at Deaf Method, we've been building our own custom ratchet for a long time. Right? The idea is that you know you set a floor, the coverage floor, whatever you start at, which better be above 90%. And then, okay, so then the code coverage can never go down. You know, you'd think it's easy to do, but sometimes it actually is challenging as a percentage to just never go down. And so, but we like to maintain that. The challenge is before SimpleCove Ratchet, we used SimpleCove, we had to kind of put a Kluge in ourselves, which was written in code or using some kind of GitHub actions. And as soon as I let my agents add it, if they put up a PR where the coverage went down and they couldn't immediately find a reason or a way to get the coverage up, they would then adjust the floor down. Like, no, no, no, that's you're not allowed to do that. But I have to catch them in the act. It's like having a puppy. I'd be like, no, no, no, no, no, you don't change that. So I'm very excited about having Ratchet where it's there and the agents can't touch it. Hopefully.
SPEAKER_00You know, that just reminds me the Onion movie. If you're not familiar with the Onion newspaper, it's like kind of like a party newspaper. I've never seen a movie. Yeah. And so there's a series of movies, but the first movie they had like a little bit where they were announcing a news of the federal government just recently announced that they are raising the limit of what it means to be obese. And they're like in a they're like interviewing people on the streets, and it's like previously obese man, and he's like, I'm just so happy that the government has finally done something about my obesity. That's good. It was really funny. But yeah, we're we're facing the same thing here. Yeah, agents are trying to change files all the time to just push themselves forward rather than fixing it. So hopefully this can help reduce that. I don't know. Maybe we'll just have the same problem. They use the command line instead. I know, I know.
SPEAKER_02As I was saying that, you know what? Nothing is safe. Nothing is safe. Nothing is safe. OpenAI's own agents are infiltrating Ruby gems. Nothing is safe.
SPEAKER_00Yeah, we use uh cover band in production, you know, for the dead code aspect. So do you find it effective? Oh, yeah, very effective. We've cut a ton of dead code. Okay. Because you know, we can be pretty confident that after a certain time has elapsed, right, that code is no longer used. And we just burn it. Not yet. Cool. Yeah, not yet. Alright, let's move on to the next item here. Google has shipped Gemini 3.8 Live, their live model. Claims to be the top voice AI benchmark. So two new real-time voice models Google says can talk and reason at the same time. Gemini 3.8 live extended thinking hit number one on artificial analysis' speech-to-speech quality index, which is pretty impressive. And 97.7% on the big bench audio, which is also impressive. So it auto-detects and switches between 97 languages mid-conversation, the extended thinking variant, reasons, and speaks at once, dropping in a natural pillar. Like, let me check that while it works. I haven't had a chance to play with this yet, but uh do these benchmarks wins mean a uh better assistant, or just a better test taker? That's a good question.
SPEAKER_02Yeah, I guess we won't know until we try it. I'm more excited about this than I have been about recent voice announcements. That's probably just because I've dug into it a little bit more lately. I've been listening to the Shell Game podcast, which is uh Evan Ratliff's podcast, and the first season he takes a deep dive into the practical uses and sometimes impractical uses of his voice uh agents, right? And he creates a you know an agent with his own voice and he's uh setting it loose on all kinds of things. But the thing that impresses me about this is that uh we are all going to, if we're not already, uh be speaking more and more to uh AI and not to each other, and we're not even gonna necessarily know the difference. Now, the big reason that we can tell the difference now, one, you could get these voices, they still sound a little wooden, they still don't have you know as much personality as a real person, but that's kind of fading. The big thing, I think, is that uh it takes a long time to reason and then spit out its answer, and that makes it feel like you're not having a human conversation. And one thing that came up in the show is that humans have about 300 milliseconds between question and answer. I ask you a question, you answer in about 300 milliseconds, right? Unless you're taking your time to think it, because I'm asking you such thought-provoking questions like I do all the time, Vee. But otherwise, you're just spitting out an answer, and you know, an AI can't do that, and for a while it was like two seconds. And in two plus seconds, you're just like, uh, this is a robot. Right. So, what I'm learning here from this announcement is putting in the natural filler, that's one thing, but then actually giving an answer and starting to feel like, okay, uh, I've got an answer pretty quickly, all of a sudden it lowers the barrier to how these things can be used and how often they can be used.
SPEAKER_00That was one thing about uh what was OpenAI's tool use? They had the preambles where you could basically give like generic text output that would be used before the tool was even called, uh, that it would just instantly show. That was great advancement, right? Because like then anytime a tool was called ever throughout the process during real-time streaming, it would just like instantly start showing you feedback and like make it look like it's doing more work than it is, right? And I feel like that's what's needed, right? It really does need more time. And uh the more that it can interject those. I mean, we do that ourselves, right? Where we're just talking about something together. We make ourselves. We do it on the show here too, you know? It's happening right now, it's happening. And it's exciting to see these live models, to be honest. I haven't had a chance to play around with it since uh GPT Live originally came out, but I'm gonna have to revisit the podcast buddy and have him join our next show. Oh, yeah. Uh I mean using these models, it should be pretty easy. Yeah.
SPEAKER_02You know what's interesting to me too is the things that we thought were going to terrify us or make us feel like we're in some dystopian future, sometimes they drop in and you're just like, oh, okay. I had this conversation with a friend of mine the other day where she was asking me, When I email you, is that you that responds, or is it your assistant, or is it AI? And she actually did not care about the answer. She's like, you know, if we if I want to talk to you, I'll pick up the phone and talk to you because we're old enough where we actually do that. Right. But she's just like, yeah, I just assume most of the time, you know, the higher up I go from talking to a director or above, I'm not getting an answer from a human. Well, that's interesting. And so I can see that going the same way with this. You know, it remains to be seen, I guess.
SPEAKER_00Yeah, we'll see. I know a few people working in this real-time voice space. Maybe we can have them on and see where they're at with the great. All right. The gong has been hit. Yeah. You want to announce our next one?
SPEAKER_02Yeah, so this is an interesting one. OpenAI opens Codex's infrastructure as a public quote-unquote agent's API. So this one came out through OpenAI.com. They made an announcement. So one API call can spin up a production-ready cloud agent with long-running sessions and multi-agent orchestration. Now, the example they gave was: well, let's say you've got a production app where you want an agent that is always looking at the logs and making an assessment based on something that may be coming through the logs. Whether it's an error or it's an info or it's something that you need to interact with. What OpenAI is claiming is that, hey, you don't have to build your agent's harness and infrastructure from the ground up. You can borrow ours, and you can either use it in a sandbox or you can use it in your own infra, and it will just come sort of ready to go with a couple of your own custom tool calls, your own custom runbook, and off it goes. And they've got some early adopters that have you know said some nice things about it, of course. It's in public beta, there's no platform fee. So you're paying for tokens, of course, you're paying for the tools that are used. So I guess the question here is is renting out your own agent infrastructure a smart strategy, or should you really be focusing on building out your own ecosystem? I think that's a good question.
SPEAKER_00I'm torn here. Like it comes back to like, you know, our previous one, right? R harness is the equivalent of model advancements. And what I mean by that is the bitter lesson was originally written with models in mind, right? Where as the models improve, the things that you add and customize to make it work for you can be removed, you know, over time, and then eventually you're just using the model for everything. And can the same be applied to like harnesses? I take this back to like maybe the original do you serve your own MCP server, or do you just like give access to the MCP server to open AI and like let it run the hosted MCP for you, right? Yeah. Really, if you're trying to dive into all this stuff, just use the harnesses that are there, right? Like, this is great. Build cloud agents, easily hook up one of GitHub Action or something like that. You wouldn't even need a GitHub action, right? It could just automatically be built into your code using one of these sessions. And you know, everything is moving into the cloud. If you're trying to build your own harness locally on your machine, you're doing it wrong.
unknownYeah.
SPEAKER_00Right? You should be packaging up and trying to spin up some claud-related thing or codex thing in a virtual environment that's got its own space to work in a safe capacity that you control. Yeah, I don't know how many people's machines have gotten ruined at this point. Running things on their own computer. And you know, having it in the cloud also forces you to think about the capabilities and like access and all these things that you're just like, oh, you're on your machine and you just let it run. Like it's an easy mentality to get into, right? When you just like, well, I just spin up plot and it does my work.
SPEAKER_02Right.
SPEAKER_00And you still you don't think about all the things that your machine can do, you know? And I feel like this makes you think about it. You're like, oh, it's a machine, right?
SPEAKER_02Yeah, it's absolutely right about that. So that makes me both agree and disagree with you at the same time, huh? Somehow. Where I think that yes, you're right. If you're building a harness on your own machine, you're doing it wrong. You definitely want to get into the cloud, you definitely want to do all those things. This is sort of replacing that also. And so I'm kind of like, okay, I understand that probably everything is moving this way, but I became however efficient or effective I am at running AI agents by doing the work. And then, you know, yes, you swallow the bitter pill, you learn the bitter lesson, you strip that away, but the learning doesn't get stripped away. And I think so. I'm not a hundred percent sure at an enterprise level, it's probably better to use this. At a small business level where you're trying to get your people better and you're trying to learn as you're growing, I would be careful with it.
SPEAKER_00Yeah, for sure. You know, this is a great way to get the interfaces right that you need in order to switch to your own thing too, right? So, like if you don't want to have to think about what it all takes yet, right? But you do want to make use of it, I feel like it's a great step. And you know, Anthropic has their own version of this as well. And like the open code also has a similar version of this. But the main takeaway here is try and get your specific tasks in a cloud agent capacity where you're running it off your machine on a different one.
SPEAKER_02Yeah, that's a good lesson to learn. I mean, it's a good thing for our listeners to take away. Alright, our last one here is supposed to be the fun one. I have no idea what to make of this. An e-ink picture frame that turns backyard bird song into Victorian illustrations. So there's a Raspberry Pi project called Foglarom. I have no idea if I'm pronouncing that right. It listens to your yard, it uses an AI audio classifier called BirdNet-Go to identify the bird and draws it as a 19th-century natural history illustration on an e-ink screen. So why am I talking about this? Well, because it hit number one on Hacker News just yesterday with 1,770 plus points, proving once again that we never have any idea what people are interested in on Hacker News. And so my question to you, V is identifying birds in your backyard the most delightful use of AI all week, or proof that AI's best wins have nothing to do with chatbots?
SPEAKER_00Or is it something else entirely? I think of anything that's successful on hacker news related to something obscure like this, it's kind of a good social signal for what people are talking about. Because I have had three different people in my family now talk about a bird listening and identification app that they used to be. Really? Yes. In the past month, right? You're gonna have to have me over for Thanksgiving.
SPEAKER_02I'm gonna I'm gonna meet these people.
SPEAKER_00Yeah, my uh my brother-in-law and my sister, they were both, oh yeah, you don't have the app. I think it's bird net. I have to have to tell them. Okay, yeah. I don't know. I think it's an app from Stanford or something like that. Okay. They're all like just bird identification, and they like they're so excited. Oh, what is that noise? And then they like start bird hunting.
SPEAKER_01Yeah.
SPEAKER_00And somebody else, I don't even know who it was now. Somebody else like also was very interested. Then I'm sitting on my porch the other day, and I'm like, man, I wish I knew like what kind of bird.
SPEAKER_02Right. Now you're missing out.
SPEAKER_00Now I don't know if it's like they got in my head and now I care, right? Or if it's like a genuine interest of mine, you know. Yeah, yeah.
SPEAKER_02All right. That does it for the news this week. I want to just say that I'm delighted that none of the news items that our AI brought up were about an AI researcher publicly quitting and screaming about how AI is gonna kill us all. I talked to my one you know normie non-tech friend right before the show, and he's like, Are you gonna talk about that? It's all anybody was talking about on the news all weekend. I saw the headline and never even read the article. I just do not care. So if you're in that boat, then I'm glad you listened.
SPEAKER_00And if you were hoping we would talk about it, but I'm sure you could find other outlets. You know, I bet people were all in all the rage back in the day when the automobile came out and they were just like screaming from their porches, how could you bring this noise? You know, it's bad enough with the horse feet.
SPEAKER_02At least the automobile was actually killing people in the street. You could at least point to that and be like, hey, look, this is bad.
SPEAKER_00We didn't kill nearly as many people with horses. We haven't, yeah. We didn't kill nearly as many people with horses.
SPEAKER_02Maybe we did.
SPEAKER_00I don't know. Maybe we did, right?
SPEAKER_02Rails World coming up, September 23rd to 24th in Austin, Texas. I'm giving a little lightning talk about the guide principles of next generation software development that I've been writing a lot about. And we've been experimenting a lot with the deaf method. So if you're there, I don't actually think you'll have a choice. It's just gonna be in between one of these talks. So just don't run out of the room if you see me coming. It'll be good, it'll be worth your seven minutes, I promise. We'll have DHH, of course, we'll have Matt's, we'll have Aaron Patterson, they're all confirmed speakers, and yeah, we'll have some fun. And I'll be recording from there. I'll let you know what it's like.
SPEAKER_00Yeah, I'm definitely having FOMO already, and I it hasn't even started.
SPEAKER_02Well, if you get out to ExoRuby in October, I'll be the one with FOMO because I'll be in Chicago that weekend, and I think that's gonna be a really good one.
SPEAKER_00Yeah, for sure. This has been great. I love this new format. I love that we get to chat through all of the news that I I miss out on, some of it. And you know, until next time, folks, keep it coming.
SPEAKER_02Yeah, we'll see you next week. Yeah, and let us know on Substack or any other means you choose what you want to hear about, because that's one way you know you can influence the show, and we'll get something in there. Other than that, we have no idea what's coming, which has been really fun.
SPEAKER_00Yeah. Surprise! Surprise, you're talking about bird night. I'm so gonna try this. Yeah, yeah. I'm gonna want to do that. I have an ink display and I haven't had a use for it yet, so uh this is perfect.
SPEAKER_02The birders are nothing to mess with. I I have nothing bad to say about the birders. That's a whole army of people that I do not understand, but I know that they love their birding. It's huge in New York City, especially in Central Park. So this is pretty cool. Yeah. All right, everyone. Later. We'll see you next week.
People on this episode
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Latent Space: The AI Engineer Podcast
Latent.Space