The Ruby AI Podcast
The Ruby AI Podcast explores the intersection of Ruby programming and artificial intelligence, featuring expert discussions, innovative projects, and practical insights. Join us as we interview industry leaders and developers to uncover how Ruby is shaping the future of AI.
The Ruby AI Podcast
The TLDR of AI Dev: Real Workflows with Justin Searls
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
In this episode of the Ruby AI Podcast, co-hosts Valentino Stoll and Joe Leo engage in a lively discussion with guest Justin Searls. They explore the evolving landscape of software development with agentic AI tools, comparing traditional agile methodologies with emerging AI-driven practices. Justin Searls his experiences with refactoring and the challenges of integrating AI tools into development workflows. The conversation touches on the suitability of AI in coding, philosophical perspectives on reinforcing proper software practices, and the future potential of these technologies. Justin also provides valuable insights on configuring AI tools for better productivity and discusses his personal coping strategies with the frustrations of modern AI capabilities.
00:00 Introduction and Hosts Banter
00:30 Guest Introduction: Justin Searls
03:13 Justin's Career and Conference Talks
07:52 The Evolution of Agile and Development Practices
16:07 Challenges with AI and Iterative Development
27:47 Recalibrating Development Processes
28:00 Adoption of Pivotal Labs' Methods
28:28 Continuous Integration and Testing
29:21 AI in Development: Current State and Challenges
30:16 The Role of AI Agents in Development
32:17 Frustrations with AI Tools
35:03 Philosophical Reflections on AI in Development
36:16 Generative vs. Subtractive AI
37:06 The Future of AI in Software Development
39:27 Balancing Coding Enjoyment and Productivity
44:02 Capability vs. Suitability in AI Tools
46:35 Prompt Engineering Tips and Tricks
52:39 Closing Thoughts and Plugs
Introduction and Hosts Banter
SPEAKER_01Hey everybody, we are live. This is the Ruby AI podcast. I'm your co-host today, Valentino Stoll, and we have our other co-host, Joe Leo, here. Joe.
SPEAKER_02Hey everybody. I would like to point out for anybody that is watching that I actually am winning the shirt war. Valentino, as you can see, now is a collared shirt. We're actually moving toward a more dressed-up version of the podcast, which is good. Get the Riff Raff
Guest Introduction: Justin Searls
SPEAKER_02out of here. We also have our t-shirt wearing friend Justin Searles here today. Justin, thanks for coming on the show. Yeah, I'm embracing my Riff Raff status. Yeah.
SPEAKER_00Thank you, sir. You should. I have been cycling through tri-blend t-shirts for 15 years now. So why mess with success?
SPEAKER_02You know, this whole thing is really just me trying to overcompensate for the fact that I don't have cool programming t-shirts, or I have very few of them. And Valentina seems to have an endless supply. So the only way I can get back at him is to try to throw catty insults here and there.
SPEAKER_00When I started my career as a consultant at like a big four style accounting firm, and they had a pretty strict dress code when I was at client sites. I don't know if things are still this way, but like I had to wear pleated khakis. And my wife hated the like flat front khakis were okay, but the pleated khakis were absolutely unacceptable to her. I could go downstairs and ask her right now. She still doesn't believe that that was like a dress code thing. Uh somewhere rattling around her head, even though I like, you know, I wear jeans or shorts 350 days a year. She still thinks that secretly I like really like plaited khaki. I like this.
SPEAKER_02I think uh here. I think having her as a guest on the show sometime to talk about this, I think would be worthwhile, be time well spent. Sure thing.
SPEAKER_01It makes me wonder what a dress code would look like for AI agents as well.
SPEAKER_02Well, something along the lines of Justin Bowen's active agent little brand or logo is what I'm thinking of right now. You know, a guy with an agent's hat and an overcoat.
SPEAKER_00Given what we've learned about how agents are, you know, you can prime them with context. If you were to start with, hey, you're clean shaven, you're all gussied up, you're wearing a really nice suit, that would probably put spin on the ball. And I wouldn't be surprised to have some researcher come out and say, hey, when you prompted this way up front, like you get 20% better performance on such and such a benchmark. Yeah, absolutely. Yes, WE is 3% higher.
SPEAKER_01Yeah. Oh, I hope that doesn't translate into the real world and we all have to start wearing pleated khakis.
SPEAKER_02Well, when they start getting some more sentience and they can see what we're wearing, that'll be where the rubber hits the road. Because I don't want to show up raggedy to my agent. To my agent meeting.
SPEAKER_01Yeah, you know, performance will improve if you're dressed nice, right? Yeah. That's how I always think about it.
SPEAKER_02But we digress quite a lot already.
Justin's Career and Conference Talks
SPEAKER_02But uh Justin, it's great to have you on the show. And we were talking just a little bit earlier about some of the stuff that you've been writing about. And actually, before I even get to that, so you've done a lot of writing in your career, you've done a lot of speaking in your career, but I saw something recently that said that you were kind of at an anniversary of your last conference talk. So are you not giving conference talks anymore? Did you say that that was your going to be your final talk?
SPEAKER_00Yeah, so if you go back and uh watch the tape, at Rails World 24 uh in Toronto, I was giving a presentation, actually speaking about my wife, about the application I built her called Build With Becky, which is a strength training app. You subscribe, and every month she designs a new program, and then that's delivered to you. It has like a workout player. It punches, like for being a one-woman operation, like way above its weight in terms of like UX and the care that went into it, because it was really about designing something for her that would fit not just her business, but the lifestyle that she wanted to be able to maintain while also being able to run such a business. And that project just because it was a ground-up Rails 7 baseline modern Rails application, it provided for me tremendous opportunity to kind of take notes along the way about like all the things that if you're a new Rails developer or an experienced Rails direct developer, like here's just like I don't know, a dozen different areas of advice for like how to do this well. And because it was for just us, I wasn't like I had been in my consulting career, constrained by what other people wanted to do, constrained by institutional stuff or like all these integrations or all the things that kind of weigh you down. Because I had just sort of this open canvas, I decided to use that to build a talk, and that talk would just be like, hey, here's a lot of great ways to do things. This is maybe where you'd look next in terms of how to actually build a product on top of just vanilla omakase rails. And all of that was a lot of work. It was like a 10-month project to like build the system and it kind of concurrently started to put together the bones of that presentation, and it just felt like a good capstone on what had been a career of speaking as a consultant who mostly went to speak to try to drum up sales for the consultancy and maybe attract candidates. But because I stepped back from that role actively at Test Double, now it was sort of just like a thing that I was still doing, even though it takes me like a month of my life to prep a talk. And then, of course, I'm a real high anxiety individual. And between the travel and like a month and a half, two months of my like existence just gets erased off the board whenever I do one of these things. And family gets frustrated with me, and I just decided, you know what, like I gotta stop at some point. And you know, it was an emotional decision because it's been a big part of my identity and it was a huge part of my career. But I think for me, when I look back at all of the conferences that I've been to, the people that I've looked up to, the speakers who inspired me, something that always disappointed me was that at some point or other they just kind of stopped showing up. Like none of them proactively announced, hey, I'm done now. And I decided because I was in the first day of the conference, you know, I'm just gonna say it affirmatively, both to hold myself accountable, but also to give people a chance to say goodbye. And I had a lot of really, really, I think, impactful conversations following that, and I have since. It feels much more tidy now, because like you could have the whole the volume set of all the talks that I've ever given. You can watch them beginning to end and hopefully appreciate some kind of arc in there as you watch my hairline slowly recede. But honestly, it felt like a good bookend on where I'd originally started.
SPEAKER_02You know, I have to say that well, I appreciate all of the thought that you put into it because it probably falls by the wayside. I mean, for those who have not given a talk, or even if you have, my experience mirrors yours almost exactly. I get very nervous before a talk, and I spend a lot of time preparing. And I've actually gotten back into it after a long hiatus where I think I just kind of forgot, forgot how anxiety-producing and and how much time that absorbs. On the other side of it, you know, I do love it. I love being up there and I love talking to folks. But yeah, the commitment is real and the anxiety is real, and and so is the travel and the time out of your schedule and time away from family. And yeah, you you want to get return on that investment. And if you're not, then it's the right thing to do to say, okay, that's
The Evolution of Agile and Development Practices
SPEAKER_02it for me.
SPEAKER_00There's a certain amount too where it's just like the world changed after COVID. I believe you guys are involved in like Artificial Ruby in the New York area. Like there are new meetups coming around, there are conferences still, but there's far fewer than there used to be. Yes. And it's easy for me to like as a cish white guy who's relatively privileged and already kind of made it, insofar as I need to in this community. I used to think about the fact that when I was trying to break in, my first Ruby conference talk was in 2011, Rocky Mountain Ruby. Marty Hott took a chance on me. And I remember trying and trying and trying to get into any of these conferences, and I remember being frustrated because it was almost like this roving like circus, as you'd see like the same 10 speakers at every single regional Ruby conference all give the same talks. I won't name names, but like people around at the time would just, you know, they were just sort of like a traveling roadshow, which is all well and good, but it just didn't create a lot of opportunities for new speakers to kind of come up and make a name for themselves. And I think that in the post-COVID era where there's just way fewer cumulative slots to be able to do that for oneself. Me kind of just continuing to do this song and dance as I increasingly either find myself repeating myself or I'm just elevated into like the clouds of thought leader land where I just become a talking head. Like we've also seen some of people who I previously had a great deal of respect for out there who gave talks a long time ago and had some good insights, and now they just didn't never stop talking and may regret that decision. So that was for me, I think, a big part of it too, is like make space for other people. You know, I said my piece, I'm finding the I don't want to just keep going on stage and play the hits forever. Yeah, I get that too.
SPEAKER_01So I will say one of my favorite talks of yours was make Ruby great again at the uh keep Ruby Weird conference in 2016.
SPEAKER_00That is the only time I ever wore a suit on stage that was in Austin, and I had a whole like Trump shtick, and uh the election was like eight days later, and boy, that aged like melt. I remember everyone had a great time with it, and then we're like, oh god, I failed to think about how I would feel about that 11 days later. Most of us did. Most of us did. It's on a curve.
SPEAKER_01Honestly, that was that was fantastic. You shouldn't be ashamed of that at all.
SPEAKER_00It was excellent satire until it was just really, really painful. So thank you for that.
SPEAKER_02Yeah. So you mentioned your kind of shifting roles, maybe how you relate to engineering, certainly how you relate to test double. I'm always curious when I come across any other leader of a consultancy, because of course I lead a consultancy. For better or worse, most consultancies like test double, like Def Method, get labeled boutique consultancies, which usually means in some kind of hazy way that we care a lot about agile software development and test-driven development and XP and pooter and you know, all of those kind of buzzwords. We take them really seriously and we iterate in our approach and all of that stuff. And uh what I'm curious about is, and you know, I'm seeing this as maybe a dividing point in the community is uh, you know, how that uh translates, or if it translates, into today's development with generative tools. And I'm curious what you think about that. And I've written a lot and we'd like to get into uh your assessment of who the winners and losers of this kind of divide are going to be, but I'm curious just up front, what happens with uh all of the good design test first development that we schooled ourselves with? Now we've got a bunch of agents.
SPEAKER_00I mean, we can talk about consultancies and services-based organizations just like as such, if we want. But to your actual question, which is really like, what do we do with this basket of practices and norms or workflows that became memefied or named and that we as consultants primulagate to our clients or attract clients through us putting out a signal, hey, we we're like this, we believe these things and then attract clients who resonate, who believe those same things, and which I think both of our companies have done a lot of. In this moment, if you're really, really fixed to that practice or to that workflow, you know, some developers love it. They love putting in headphones, sitting down, writing focused code this very particular way that they've been honing as a craft for years, just like woodworking or something. And if you're in love with the way that you do things, then there's a certain inflexibility, right? You've ossified to a certain extent because you've optimized. Just like if you write code that is hyper-optimized, what you're doing there is you're reducing the flexibility, you're adding like additional inflexibility or rigidity to that code in order to increase its performance, typically, is how like most optimizations show up. You look at them years later and be like, boy, that looks really arcane. And it's like, oh, well, if you just rub wax paper on the slide, you can go down 10% faster. And if that's you, if you're just like you love coding your particular way, then right now, with an upheaval in terms of what is the most productive way to produce software that is being, you know, requested by an employer or by a client, with that being up-ended, your favorite way of writing code is under threat. Now, the flip side is so some people get to those practices that particular way. And to a certain extent, I did. But the flip side is like if you arrive at those practices, not because you're in love with perfectly organized code, for example, or test-driven development or any particular style, if instead you wind up there because as patterns, you find that they're just you were shopping for the most effective way to write software, that was it. And if something better comes along, then you'll embrace that. And if this is the something better, then you're happy to discard the practices that you don't have to do anymore. It's like, okay, cool, I travel light, right? That was like a common expression of a value in that era, in the agile era, like travel light, you know, like start every project. You know, a former colleague of mine called his process null process. You'd start every project with no process whatsoever and only add things as they proved value, forcing himself to kind of purge himself of all of these presumptions going into the new project. I used to hear it sometimes as strong convictions loosely held. Yep. That was another popular meme. And so the community would attract both types of people. Yes. Like, you know, I hired somebody once at Test Double, and I'm amazed this didn't happen more often, that was just extremely fixed on like, we're gonna do exactly this thing. Finally, I get to work for Justin because he's like completely anal retentive, like I am, and likes it, clearly likes it exactly this way. And then I had to like break the news to him. I'm sorry, like when you work for clients, delighting them is the job, not writing code your favorite way. I expect to get paid a high salary to write code my favorite way, a spirit of entitlement, and that's not gonna fly anymore. But if instead you're like, hey, you know, like I practiced a bunch of TDD, I did all these mocking libraries, it helped me express this particular way to organize code. Now you've got to reevaluate, hey, in the era of agentic AI, is that the right level of abstraction to even care about? And maybe the answer is no, and you got to be comfortable with that. And so, all that to say, which disposition you find yourself in will almost certainly determine how you feel about these agentic coding tools. And that's why I think we're seeing a lot of people in the quote unquote gives a shit about how they program community, where half are like all in and half are extremely intransigent and if anything, like opposed to using these AI tools.
SPEAKER_01You mentioned Agile. That was something that Rails early on adopted. And maybe what made it successful earlier on, there are a number of reasons. I feel like that work ethic definitely helped the startup world accelerate more efficiently, but also like manage itself once it launched, like adopting those patterns and having some kind of rigor attached to that in a systematic way. And so I'm curious, like, we're kind of seeing like a devolving of business practice as we know it, as far as software is involved, right? Do you see anything replacing it in terms of patterns or systems that are worthwhile, or is it still very much like higher test level and we'll figure it out with you?
SPEAKER_00Well, Agile had its moment where it was for the people who are dialed in, and I think this has more to do with like how people who are attracted to new movements or new ways of working or new communities, or you know, sometimes that's like a process thing, like Agile, and sometimes it's a language thing. I'm not surprised to see that like a lot of people who joined the Node.js community in the early days and wrote a lot of the first libraries and packages for NPM got bored after a few years once it started maturing and then went off to Golang or went off to Rust because that was the moment that they most enjoy in that adoption curve. And so anyone who joined something early, like did in Agile between 2000 and 2008, they are a different kind of bird than the ones who joined Late. And so Agile after 2010 was a lot of like now it's big institutions and enterprises, like it was such a catchy word. It was actually the marketing was too good. We often joked like lowercase a to uppercase a agile. When it got corporate, when like you know, the scrum guys start selling certifications and stuff, it's not about the process, it's about the beliefs that inform the process, right? And so, like, it became about the process, it became about adherence to this very strict and ossified rule set. And so, in that sense, we stopped saying we were agile because like that became associated with the bad way. If anything, you know, and Valentina, to your point, I kind of feel like in our community, in the like open source adjacent give a shit about how they program community, we stopped talking about agile, and everyone kind of it devolved in the early 2010s. It devolved with GitHub, really. It was like now we all just kind of do the GitHub issues thing, and we don't really think very hard about how we do our work. And you either have been doing this long enough that you picked up those skills, or you're just kind of following the issue pull request, and you have no real additional rigor or process above and beyond that, other than what is kind of implied by just a sort of a Git workflow and a GitHub specific kind of layered on abstraction on top of that. And so we've sort of been in this relatively low process state for a super long time now. And I suspect that when we talk about how agents affect that or what are we doing now, it's an opportunity for us to reinspect. That's why I keep finding myself having conversations about Agile. Because I, if anything, I think it's like because we're at the beginning of a brand new curve, it's an opportunity to kind of like dig up a lot of that timeless wisdom that has been discarded from previous iterations of the same sort of technological upheaval.
SPEAKER_02You know, I think that's so interesting what you said about Agile and about how moving on to GitHub, all of a sudden we started without really talking about it, we started adopting a much lower process or much less process. I remember having this conversation with one of my principal engineers in probably 2021, but I could be wrong about the exact year. Where I said, okay, well, so are you guys doing like your agile development lifecycle or something like that? And he was just like, you know, I don't know, we just call it iterations. It's we're with the client. At some point we got to have a meeting and show them what we're working on. At some point we have a meeting and we talk about what we're gonna do the next week. And uh, you know, the rest of the time we just code and we pair on our stuff or we review PRs. And I was like, oh, right, that's a good thing that is evolved because they're doing enough to please the client and we're a consultancy and make effective use of their time. And there's really nobody arguing that things need to be demonstrably better. Things are efficient and they're just moving. So I think that makes a lot of sense to think, okay, well, we actually already are changing. We actually started a long time ago. I am curious. So I saw this kind of couple of posts between you and Jared Norman, and I regret I wasn't able to go and check out the podcast with the two of you. But you know, a couple of months ago, and I think this dovetails with what we were just talking about, there was a discussion about entry points, right? Where we start. Where does a programmer start? And I liked what what both of you had to say, and sort of it seemed like Jared was, you know, accepted what you were saying as a refinement of what he was saying. But what I wanted to see more of was the part where Jared says, Hey, actually, if you uh prompt your agent in such a way that it starts to build iteratively, then you get better results than if it just went off and started doing its own thing, right? And building multiple files, multiple test files. And I was curious if you have had a similar experience, if uh there is any merit to iterative development with an agent, or if that too really needs to go by the wayside.
SPEAKER_00I think uh definitions are probably important. Uh in terms of what even is an iteration. Iteration. It's a loaded term because it means lots of different things. Okay, that's fair. Based on what like level of organization we're talking about, if we're talking about an enterprise, what's an iteration? Well, that might be like a release, right? Like, and so maybe that's a quarterly thing, because then we're communicating with other departments. If you're talking about like a team-wide thing, maybe that's weekly, and maybe we do a weekly, like, you know, internal release. But that's like, okay, so what are we doing this week? What did we do last week? You know, like what do we have to demo and deploy the traditional agile iteration kind of mindset? And at like a personal level, like as an individual or as a pair of programmers, maybe an iteration in that sense. Like, what are we iterating out? Like, what's the loop? Is the real question. Yes. Well, maybe the loop is I pick up a feature card, I do the work, I get it tested, whatever that looks like, I QA it, and then I merge it in, right? Even within that, right? It's like I think of all of these, whatever term you choose to use, and just iteration is a good one because it's what we're doing here. I think of it as concentric circles all the way down. And so traditionally, prior to having coding agents where the atomic unit of a programmer was the human typing code, the smallest in that concentric circle is like, okay, code isn't doing this, code's gotta do this. Maybe it means I write a failing test and then I run, and then I write some more code, I change the message. And then like the wider loop is I get the test to pass. And the wider loop is I get to like, you know, kind of bundle up that handful of requirements expressed as tests into some sort of functionality. Now, what we're dealing with here, and this is why I challenge the definition of iteration, is we're figuring out, like, oh, it turns out like we can zoom in the microscope a layer lower and have one programmer running one to many agents. And what's an iteration in that context? And so for me, the general guiding principle that I've had is like, what makes a good feedback loop for me as a human? How can I apply that same wisdom at that lower level? And the current tools, I you know, I do not assume that anthropic and open AI are chak a block with Midwestern software developers who came up out of the agile software community. There's no indication that that's where they're from whatsoever. It's much more move fast and break things startup sort of stuff. And you see it in their releases and the quality of the software, but also like one assumes the design of these CLI tools, the system prompts, how like testing's an afterthought. And I'm not here to harp on testing, but it's just like the way that these tools want to work out of the box, where they kind of like big bang plan a big old fucking task list and then just kind of work through it all at once, and like whether it's automatically just kind of gonna force itself all the way to the end state. It's not collaborative, it's not like pairing, it's not working in small chunks like I might design it. I find myself trying to coerce these tools into a adopting things like working, identifying really small chunks, breaking the problem down into something so small that you can't help but succeed. Stopping early, like hitting the brakes. So the Toyota production system was like a big part of its introduction in Japan, of course, but like also its impact on lean and systems thinking. It's like there should be a big fucking button somewhere in the room. When anything goes wrong in the production line, we hit stop and then we swarm together and we solve it. And that is also lacking. So here I am just like holding down, like, you know, like I had to monitor its chain of thought in my terminal and like mash on that escape key to get it to stop as soon as I see it go off the rails. Because none of these things are very good at being like, hmm, my confidence just dropped from 0.7 to 0.6. Maybe I should stop and ask my human. That's lacking, right? And you can try to like construct stuff on top of that in terms of patterns, like get one agent to write tests and the other one to do stuff. But frankly, like I find myself in this frustrated state where all I can really do is like beat the monolithic procedural process out of these things on a day-to-day basis. And it's not fun, it's frustrating. No amount of prompt engineering has gotten me to a state where it'll actually default to working the way that I would, which is in small incremental steps that are each validated with real use. You can get close and there's moments of brilliance, but then there's going to be days where, like right before I joined this call, I was telling you before we started recording, I was kind of mad because I had just gone 15 rounds with codex after it had done everything else right, just to change the margin on one little tailwind thing. And I was like, just blow this out a little bit, just make it a little bit wider. And it messed up four times, and then it made it go in edge to edge with its container. I was like, no, just bring that in a little bit, and then it suddenly it was like slimmer than ever, and now there's more margin than ever. And so I gave it six more chances, and each time it would just get slimmer and slimmer and slimmer because it forgot how negative margin works in CSS, and then there's no way to roll back, right? And so, like this lack of safety, this lack of consistency and process, all the things that made my parasympathetic nervous system feel good about programming, where I was like, Yes, I got the safety of being able to move incrementally and always safely roll back, all that stuff's like out of the goddamn window now. And all of the tools that we build to constrain these agents to try to get to the outcome that we want, I find them all to be lacking. And if I get them to work, then they break the next time that there's a model update.
SPEAKER_02V, you say it looks like you were smiling and laughing. Yeah, I mean identifying so many things here.
SPEAKER_01Yeah, I mean, if the ARC challenge has proved anything, it's that LLMs have zero spatial awareness. I see that every time when I'm trying to generate an image for our podcast, and it like just cuts off text on the border, and I'm like, hey, just add some space next to the border, and it just like gives me back the same image. Right. Right. There's just yeah, there's no idea about space what that means. So it's funny because I feel like we're in this recalibration state where like we've kind of like adopted, especially Railshops, you know, have adopted the Pivotal Labs way, right? Okay, we have like some kind of feature or refactoring or bug, and we break it down into like minimal steps, and then we assign those tasks and then funnel it through whatever process that we want to set up for that. And you know, GitHub adopted the same process, and it seems every other startup out there in the Rails community has adopted similar processes in different ways with continuous integration and all of this stuff. And people follow, you know, they try and follow test-driven development in some ways, some people don't like test-driven development, and they'll just have some tests that verify certain stages. We have all this contention about where to test, but really it's just like you're trying to make sure that the thing you're building is reliable. I think that's what the whole system is. So it's interesting to see kind of how different people are adopting this kind of pipelining and reassessing where the the agency of these tools fit in. So I'm curious, like, where do you see the parallels starting to evolve from that kind of business design system where we had these predefined processes for splitting up work and assigning it? I'm with you, I don't see an AI agent taking on the responsibility of like picking up a feature that's been well broken down even into very defined specifications and like running with it. But I do find that if I have an idea and I just want to see if it's worthwhile to build, I'll just throw that in the background and have something work on it and see what it produces, and then be like, oh, that wasn't worth it, but sometimes it is worth it, and so like I'm trying to also figure out kind of where these things fit in best. So I'm curious, like, where have you seen the strengths of it really going to? And where should we be like aligning as these full breadth developers, as you're talking, where do the two merge, right? And converge on on some kind of like systematic approach.
SPEAKER_00There's where we are now, and we can only talk about where are we now. All these tools at any given moment are knock on wood, the worst that they will ever be at this job. And so if you listen to this a year from now, hopefully my following statement is no longer true. And that is right now, in my experience, for the individual functions, the needs that have to be met. We've been talking about test-driven development. You mentioned Pivotal, like we're in the Rails community, we all have a certain shared context, but you could be talking to like any developer who's never heard about any of this shit. And like at the end of the day, it's still like what needs to get written, why, how is it supposed to behave in detail, how do I make sure that it's working in an automated way? All these needs are universal. And so, how do we get these needs met and like what order do we like that to happen in? Those particularities are interesting, but fundamentally, because I have a particular way that I've seen work a lot, if I had a junior developer or if I have a coding agent, I'm gonna probably direct them to kind of check these boxes in the order that I would, you know, like, okay, so like you've been asked to do a new thing. The flow chart probably starts with, do you have any clue on earth how to do that thing? And if the answer is no, then it's like, all right, well, we're gonna go do a research spike. We're gonna do something over here in a safe space and just like write a shell script or do some research and then just do a little toy app to test out whether or not this is even feasible. And through the learning of that, then we'll be able to go to an implementation and not just make a gigantic mess in the process, right? And for giving you know a spike as an example, before you step into a full-blown planning phase or whatever it is, I have found that whether I'm working with a very inexperienced developer who just doesn't have that muscle memory yet, does not have that hypervisory ability to see the forest for the trees and to kind of go through each of those individual functions and deploy them in the right moment in the right context. That's when I would have to step in and say, okay, let's do this next and this next and this next, or how about trying this? And I find myself frustrated because I can run 10 agents at a time, but none of them exhibit that hypervisory like competency or maturity. So sure, I could run 10 agents at a time, but that means I'm in the role of 10 times N being a minder to say, like, oh, do this now, do this now, do this now. And I think that the maturity level right now is, and you see it too in like AI first engineers who are either novice engineers who came to like, you know, an agent at coding, and this is like kind of like their first a lot of like nascent like bloggers and TikTokers and YouTubers who clearly aren't from the same school of thought or or level of experience of us as developers, where it's like, whoa, I got this brand new idea, guys. It's called spec driven development. And you just like write what you want and you have it write down the plan first, and then you separately tell it to do it. And I'm like, holy god, okay, sure, yes. But like none of the tools are exhibiting any level of capability at actually orchestrating a bigger process above and beyond that. You see, tools, I think you guys talked to Obi probably, he probably talked about his like rail swarm. What's this thing called? Claud swarm? Claude, yeah.
SPEAKER_01He does Claud on Rails, Claud on Rails.
SPEAKER_00I think, yeah, yeah. Which has the swarming thing where it's like personally, it doesn't make a lot of sense to me, but it is a way to do it to say, like, you're the agent in charge of models and you're the agent in charge of controllers. Like, that seems like a lot of throwing stuff over a fence and hoping that it all gets plugged in together right. But short of having some sort of equivalent substrate or superstructure, these tools are not on their own, like handling that hypervisory level very well. And so I find myself continually frustrated that I'd be like, for the 30th goddamn time, it's like, yes, and now you have made it work. So let us write a test that makes sure that it keeps working. It's like, oh yeah, good idea. I'm like, geez, like like yes, they're telling me my ideas are good. I know it's a good idea. I came up with open the browser and check it, you know, just like the last 30 times we did anything, or like your agents.md or claude md file says, is like, how to open the browser, how to log into the app. Did you check it? And the answer is, oh no. I it was like next steps, you run the tests. I'm like, Yeah, you run the test. Like, that's what you're here for. And that's been very frustrating that we're not able to provide that it can't provide for itself, but also there's just not a really great and feel free to write in Justin at Cerles.co if you've got a real lock, stock, and barrel wonderful way to force it to do these things, but like clawed hooks ain't it. There's just not a clear way to enforce or to even encourage a really, really successful. We've got reinforcement learning, we got all these things on the table, and none of that's being deployed successfully. It's really disappointing.
SPEAKER_02So that kind of brings me to maybe a philosophical point. Or maybe it's just sympathizing with those that are maybe not on board, right? Because hey, I I do it my way and I love my way, right? These people that we talked about earlier on in the show. We really have to fight the agents to get it to do the things that we really need it to do. And it makes me wonder, okay, so when we do this fighting, when I do this, like when I become a cudgel and I just bludgeon my agent into doing things with the way I believe are the right way, am I uh enabling the AI to do things properly, or am I actually standing in the way? And I guess what I mean by that is is there some other practice lurking out there? Now, I don't really know what that is because I think to me the alternative sounds like chaos, right? We just release stuff that we have no idea if it works or not. But I do kind of wonder from a, again, from a philosophical perspective of okay, so how much are we supposed to fight this thing and how much are we supposed to figure out how it works and go along with that?
SPEAKER_01You make a really great point. And what I was thinking about while you were talking about that is how it's generative AI, it's all additive, right? Yeah, and that's why we have this problem of hypervisory mechanisms, is because it's meant to just generate more. So if you just have a hypervisory layer that is using the same generative thing, you're just gonna get even more, right? That's there's no the bear style subtract, right? We almost need like I don't know what you would call the opposite of generative, but like subtractive AI, subtractive AI, right? Like that literally does like sit there in that hypervisory layer and takes away and like whittles away and reduces tokens. I don't think there is anything like that that I know of, but it does make me want to build something now that's interesting.
SPEAKER_00I will say to any of the uh doubters who nevertheless listen to a podcast with AI in the title, there's bound to be some. Yeah, sure. Look, it's gotta be frustrating because this is not only maybe not your perfect favorite way to write software, it's not only trained on data that was stolen, right? Unethically or illegally potentially, it's not only gonna disrupt job market after job market after job market, including our own, ultimately resulting in a place where maybe there are fewer total programmers who make good money, right? All of these things can simultaneously be true. You're not necessarily just being a stick in the mud if you also are looking at like somebody like me who uses these tools all the time, be like, boy, Justin seems like miserable. This does not sound like he's having fun. And that's the answer is I'm not having any fun. And the reason I'm not having fun, the reason this is such a torturous moment in time, is that I can hold these different things, these seemingly contradictory points in my head at the same time. Like, on one hand, these tools are already on net more powerful than I am as a programmer. I have seen myself get the baseline, the worst, even the most frustrating day with these tools, I'll call it a 1.3x output versus if I just coded by myself. Easily. And on the best days, maybe it's 2.3x, right? And because I know that I'm getting more output out, and it's at my level of quality, and and I define quality more extrinsically now than thinking about like that everything's named just right. Although, Valentine, to your point, I have had to learn to just like let go a little bit. Because I know that I can get more out of it, and I know that this is unlikely to get worse, right? Like these tools are only gonna get better. This is a skill and a an approach that I need to learn and I need to master just for my own purposes. When I blog and stuff, I'm just kind of sharing what I learned in real time. And if I were gainfully employed right now, I would be very worried if I was not getting on this bus and not gaining those skills because it's gonna be the new baseline, if it's not already, in terms of how we're expected to be able to both be productive and also how to like, you know, interface on teams and how to show up at work. All of that is like extremely frustrating when viewed in the context of coding can be enjoyable. You can focus, you can work, you can solve problems, you can think through things, but juggling for terminals that are kind of, you know, all taking turns, working in the background, and you're context switching constantly between the one that is trying to like, you know, overhaul your README to the one that's trying to change how you authorize against the Instagram API, to the one that's trying to like fix your CSS and another thing, and they're all changing at different times, and you've got to go switch and you're just kind of playing hot potato to keep all of the different agents fed with requirements and with feedback, and then you're constantly frustrated and you're mad at the computer for writing the code wrong. It is a much, much more frustrating and less satisfying, and just you feel hollowed out at the end of the day. And I don't feel great about that at all. Like I don't know what to do with that information. But I simultaneously I look back and I'm like, wow, this system took me half as much time to build or a third as much time to build as it otherwise would. Like, what am I supposed to do with that information? If, especially if I'm an employee or a consultant and I bill as a unit of time, right? Yeah. If I'm expected to like make home a salary, it's like if I'm gonna take two times as long just so I can have my perfect day at work every day, eventually someone's gonna sniff that out and say, hey, like do that on your own time. Right. And uh Rails, I think it was Rails Conf, I remember talking to Aaron before he was doing his keynote. You know, his whole point was maybe all this agentic coding will help us get back to what made Ruby great in the first place, which was it's just a hobbyist thing. Like I can write Ruby as my hobby and let all the drudgery, all the work-related stuff be the automated stuff. Now that's wonderful if you don't need a job, but like at the end of the day, this is just where things are going. I don't know what to tell you, right? Like, yes, I am miserable, and also this is probably the right answer.
SPEAKER_02It's interesting because, as you said, if you're not doing this for an employer, you could write this any way you want. So there's got to be something uh in it for you, right? If it is just the future or it's efficiency or something, right, or else why do it? And I'm curious, I mean, is it uh amidst all of this, I'm not having any fun right now, is there some optimism about the future?
SPEAKER_00No, I I think humanity is pretty boned and we're all here. Why do I do it? Because like while I can like coding, I can like lots of things, especially things that make me money. And while that era of my job is past me, I don't code for the experience. I code to have built things. And so, like, what I really want are like two or three large important pieces of software so I can live a better life. And if I can get them three months faster, two years faster, if I can unlock the ability to like turn my little random ideas and my backlog of app ideas and project ideas, if I can turn those around inside of a day or inside of a weekend where before that might have been a week or month-long project, that's reclaiming time for me just to live and not be behind my computer. That's why it's worth getting frustrated about. I'm willing to trade two weeks of misery in exchange for six weeks of freedom, right? That's the trade-off. And so it's not just about money, it's about are you focused on coding as an experience, problem solving, like maybe other people do crosswords and I write code every day, like the advent of code kind of thing? Or do you code so that you can build useful stuff? And I've always been in the latter camp. I'm about the outcome. I can get really into the process. And Valentino, to your point about like all the stuff I've written and everything, I've written a ton about that stuff and I care a great deal about it, but I would throw it all away in a heartbeat if I could just snap my fingers and get all the software that I want.
SPEAKER_01Yeah. Yeah, that's what I want. I just want something doing all my work in the background while I'm playing ping pong, right?
SPEAKER_02Well, at the end of the day, you know, the outcome is why we start. I've always been in love with the process. At the end of the day, it's always because I believe it's productive. And productive has some very, you know, vague meaning that's different to everybody. But to me, it's being focused on the outcome is a worthy goal. That's necessary at the end of the day.
SPEAKER_01Ultimately, that's what we're doing anyway, is just watching the output. So it makes sense that we care about the quality of it. I can't say that that's the same of people new. Do they care what the output is? If I was starting out, I would just generate it and keep generating, right? And keep producing. So I'm a little worried in that respect. At that point, hopefully things just get improved and somebody can build something better that. generates more quality.
SPEAKER_00Gary Bernhardt did a talk that I think was at like SCNA, the Software Craftsmanship North America Conference in 2011 probably. And I don't know if the name of it was this, but it was Capability Suitability. And what he did was he charted over time that in the history of computer science that we've had these momentary spikes in capability. Like suddenly we've jumped up the stack and computers can be used to write way more like express brand new functionality in ways that we never really could in a practical way before. Right. You go from assembler to C, and suddenly like the category of applications that can be written by a human expand dramatically. But with that rise in capability comes sharp edges. You know, you got pointer math. You got all these ways to fuck things up along the way. And in general the next thing to follow the next big trend is a suitability like kind of putting the bumper bowling lanes in the side of the alley to like prevent you from going off the rails. And so that's why you'd have a whole bunch of suitability adjustments like Objective C and classes to try to rein in some of the unstructured bits of C. Yeah C itself in a lot of ways you'd you'd have Java kind of the ultimate form of this of like safety scissors for programming but like still very much a C-like language at the end of the day. And what we're experiencing right now is a capability spike unlike anything that we've ever seen. Except it's not about code per se, although I'm sure at some point somebody's going to release a programming language designed for agents or at least claim it's better. And it's about how much throughput can a single human get in terms of code production. And we're seeing all the excesses of that. We're seeing the vibe coding we're seeing you know junior developers and senior developers alike just pushing up massive pull requests that are written in a day and not really reviewed very well. We're seeing all kinds of downsides of having this tremendous capability spike and where I'm at because I'm more of a suitability guy it's how I ended up starting a company with test in the name I'm trying to figure out like where are even the seams to rein these tools in to fit a process that is going to result in more predictable outcomes. And we're just not quite there yet. We've got some of the primitives we would need to build that but I find myself frustrated the suitability stage isn't coming soon enough.
SPEAKER_01Yeah so I wanted to start a new segment at the end of our show here where we share our prompt engineering or vibe engineering or whatever you want to call it our most valuable suggestions for the listeners what's working well for you like what is your go-to use of an agent or maybe just a prompt itself that you've been finding I've been doing a lot of refactoring and so I have a kind of script that I run through where I ask it to refactor as if it were experts like a panel of judges and I pick ones that I know it has a lot of training data about. So namely Sandy Metz, Martin Fowler and Kent Beck are kind of my trifecta that I source because they all have a lot of literature about refactoring. And it happens to align well with Ruby in a lot of ways. And so I do find the results are pretty remarkable. So that's kind of my prompt I've been using lately. I like that.
SPEAKER_02Well two things one I would say if you happen to be listening to this and you use AI for anything that is not writing code I would say that the act of creating GPTs this has kind of been a total dud from a business perspective from OpenAI. I nonetheless find them really useful. So I create my own GPTs and they do things like generate proposals for me or write contracts or no they don't usually write contracts but they'll do a whole variety of things that's really helped me in the business side of deaf method. I also just wanted to take the opportunity to call out a thing that I read just when I was doing some research for this episode. So Justin you've eliminated retained context from your GPTs or from your interaction with LLMs. And I thought that was really interesting because I have been frustrated lately with because you can't choose what these things remember. And so I'll ask a question and I'll be brought in some context from something that I mentioned like two weeks ago will be brought in and it's really frustrating. So I haven't tried it but I think that's something I'm going to try next.
SPEAKER_00Yeah strongly recommend turning off any chat history or personalization stuff that's like referencing other stuff because all you're going to do is waste tokens on stuff that's not relevant. I would prefer not to contract psychosis from my relationship with chat GPT and one way to do it is one way to do it is if Command N basically gives you a clean slate just like a logged out Google search or something right like I want to see what you would tell just anybody off the street. I don't want you to like put spin on the ball based on things you already know about me.
SPEAKER_03Yeah.
SPEAKER_00Because I started seeing it say hey you know based on the fact that you're buying a condo in Japan I'm like whoa wait a second we're talking about yeah yeah exactly my pro tip and this is something that it took me a few months to land at so if you use a terminal based coding agent I you might know that like so if it's a claude you have claude.md and if it's codecs it's agents dotm it looks for these files in multiple places. And the two most important ones are it's going to expect one in the repository route from which you are running the application. So if I'm working on an app called PosseParty right now at the root of that repository I have an agents.md because I'm using codec CLI. Additionally though it's going to look for a global one. So in my home directory I have like dot codex slash agents dot md. It'll load both of these things up. It'll load both the global rule set as well as the local rule set. So maybe you already knew that and if you did you like me might be wondering like what is the best stuff to put in there both at the global level and the local level and for a while because I'm working in multiple languages at once like I was like well the global one shouldn't have any Ruby specific stuff in it because this is a Ruby project but not all my projects are Ruby. And the local one should probably not have like all of my timeless wisdom that's just been piled up over time because then my other projects won't take advantage of that. Where I've landed and I think this is working well or at least it's making me feel better about it because at the end of the day this is all just a goddamn Ouija board of us kind of hoping that the system's going to listen to what we tell it to do.
SPEAKER_03Yeah.
SPEAKER_00My global one is structured as a whole bunch of bullet points. And I'm happy to share this with you even though it's long after we get off here. It's basically to say like on this team we've all agreed to always do these things. And if you can't do this thing stop and ask your pair the user and I list out always do this, this, this, this, this, this, this. And then additionally on this team we've agreed to never do these things. And if you ever find yourself really wanting to do these things stop and talk to somebody we never do this and this this is and I've got about like 20 of each of these things. And so that way whenever the agent does something I really really don't like I have a place to put that in the always or the never column. And like I said is it working I don't know but I feel better after I write it. Then the local one can be just examples. It's like example how to log into the system as an agent. Example like here's how I like my unit tests to look with all the tools and it's like I'm using Mocktail on Posse Party for isolated unit tests, but for integration tests I'm doing this and for system tests I'm doing this. And give it one good example that it can kind of use as a golden record for like when I'm doing this for the first time. And I found example based as opposed to just edicts is generally more effective because you'll basically just be anchoring that versus the norm. So you know it still does its own thing all the time. I find that one two punch to be pretty helpful.
SPEAKER_01Yeah if you're out there building libraries like Ruby gems or or something like it please put examples in the documentation strings I find any library that does have that I use context7 MCP server. It does much more reliable output for libraries that use and document that and so I've been adding that more myself. Cool well you know I appreciate you coming on Justin this has been awesome I think we've cut into some deep sections of where people should be focusing their attention on what value is getting delivered by all these things. I really love your content about those particular topics. So please keep sharing as you go because I didn't mention it before but I used to have this Twitter bot that would just run on my local machine and anytime you mentioned neat in a post it knew that you were really annoyed by something that had happened. Yeah I had it just like post thing up on my screen every once in a while like oh like look at this funny thing that he found.
SPEAKER_00Yeah cross reference that to my blood pressure or something yeah I appreciate that Valentine I I just want to like mention a couple quick things before we part ways I'm bad at plugging stuff. I'll say if you're a listener to this and you don't always already follow my stuff justin.surles.co is the place where I put basically everything and then I'm working on an application called Posse Party which I hope to release soon that is syndicating it to all the places. So if you see me on any of the social media I'm not really there. It's a robot who's just putting links back to my website on there. You can follow me there if you like the sound of my voice and you don't mind listening to me talk about stuff outside of programming I have a solo podcast called Breaking Change two to three hours an episode talk about life talk about technology games whatever's interesting to me I it's just drive time radio something to do while you wash the dishes but I also started doing shows within a show so I I have an interview show of my own called Hotfix. And the reason I raise it is that this conversation I think is pairs nicely with the one I just had with Jose Valim about how we each use agents and how he expressed that he'd be really sad if he didn't get to program right anymore in the future. So maybe you'll find something interesting to listen to over there.
SPEAKER_01Yeah excellent definitely check those out I've been loving the episodes so great deal I'm a listener. Yeah yeah thank you it is pretty great it's casual listening but also like I feel like you bring up a lot of great points and get people to talk about things maybe they don't feel comfortable talking about which is always fun for me.
SPEAKER_00Every show needs a goal and my goal is I want to get the guests to say something that could get them fired.
SPEAKER_01Uh keep the heat up we'll keep it up I appreciate it. Alright well thanks again Justin yeah I've had a great time here and hopefully we can have you back again uh another time. Yeah it's great having you on Justin thanks so much. Right on.
SPEAKER_02Well you know where to find me now.
SPEAKER_01Take it easy guys.
SPEAKER_02All right
People on this episode
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Latent Space: The AI Engineer Podcast
Latent.Space