Back in January, I received a note from a senior software engineer in Silicon Valley. He described himself as an AI skeptic who became converted after trying Claude Code for the first time. “Overnight, it changed the way I do my job,” he wrote. “It’s really, really good.”

As he explained, he no longer used a standard development environment. Instead, he “exclusively uses Claude Code” to get the job done, interacting with the tool in a terminal window and allowing it to program on his behalf.

“If I had to guess,” he concluded, “I’d say a task that would have taken me a week now takes me 2 days.”

This past winter, when I surveyed more than 300 software developers to learn how AI was transforming their jobs, the majority told a similar tale of shifting from writing their own code to instructing AI agents. The speed with which this new tool became ubiquitous in this industry was stunning.

This story matters for the rest of us because AI coding tools have emerged as the prime example of the power of AI—the first step of many more soon to come on this technology’s disruptive march through our work and our lives.

But what if the reality here is more complicated?

Last week, I received a new message from that same senior engineer who wanted to share an alarming addendum to his tale…

“I’m writing to give you an update on my current thinking about the state of AI in software engineering,” he began, “because my attitude has shifted quite a bit.”

He told me that features he generated using Claude Code ended up crashing their product on two different occasions. His boss told him that if it happened one more time, he’d be fired. “I’ve never had quality issues like this before in my career.”

The problem is that code produced by an AI agent looks reasonable, but can contain ‘hard-to-spot bugs’ that end up causing major problems. As a result, you should carefully review your agent’s output, but this is difficult. As the engineer told me, it’s “famously hard” to understand code you didn’t write yourself, so this extra step becomes “easy to just blow it off (especially when we are all trying to ‘10x’ our velocity).” Soon, systems start to break.

“The coding harnesses are useful and make life as a developer easier,” he summarized, “but they also encourage laziness.”

In response to these issues, this disillusioned engineer has returned to largely programming by hand. Here’s how he explained his current philosophy:

“Writing your own code, slowly but surely, and using LLMs for narrow or particularly annoying tasks (say like writing tests or throw-away scripts), is the best way to produce the highest quality code, since it’s the only way to properly understand it.”

Here’s the thing: he’s not alone.

I increasingly hear similar rumbles from many other people in the software industry (see, for example, ​this podcast episode​ from May). Tools like Claude Code can feel like magic, but the strategy of outsourcing all code production to AI isn’t currently sustainable.

In addition to reliability issues, it often engenders a mind-numbing workflow and an environment where junior developers will never acquire the expertise to become senior developers capable of designing complex systems.

Meanwhile, as the frontier labs reduce their subsidies on underlying computing costs, the old habit of burning through as many tokens as possible in search of workable results is proving prohibitively expensive.

From the outside, software development seemed like the poster child for AI’s potential. On the inside, it’s a mess.

This doesn’t mean that coders will abandon AI; its facility with programming languages is too valuable to ignore. But I think there’s a lot more work to be done trying to figure out how to integrate AI into this industry in a way that actually works.

This is a key point.

This last year has been exhausting. The PR departments of the frontier labs have done an excellent job convincing us that AI developments are occurring at an astounding, world-changing rate. But if you zoom out, it becomes clear that almost every “breakthrough” since last summer has concerned the narrow domains of computer code and math, which are defined by highly structured languages and come accompanied by massive amounts of specialized training data.

And yet, even in this best-case-scenario setting for AI, we’re still struggling to figure out how to actually use these tools in a way that makes sense in the long run.

This doesn’t mean that AI doesn’t work or is useless. But it does emphasize an important truth: AI is not a magic “infinity machine” that can solve all our problems, and ultimately deliver us a sense of meaning in a cold, confusing world. It’s a normal technology, and perhaps it’s time we start talking about it that way.

  • Feyd@programming.dev
    link
    fedilink
    arrow-up
    17
    ·
    edit-2
    1 day ago

    Another aspect is that “some of the more tedious tasks” already had tons of ways to be sped up but the brain-dead LLM pushers had never actually tried to optimize their workflows before and thought LLMs would make them be as good as their much more competent peers without actually learning anything.

    • Buddahriffic@lemmy.world
      link
      fedilink
      arrow-up
      1
      ·
      edit-2
      3 hours ago

      Ironically, LLMs are pretty good for helping optimize non-LLM workflows. Brainstorming a list of things that already exist is something they do well. And even if they hallucinate some entries, just move on to the next one. And having them help with one time setup or debugging specific issues is way more efficient than having a workflow that involves them regularly.

    • expr@programming.dev
      link
      fedilink
      arrow-up
      5
      ·
      8 hours ago

      Seriously. My argument against “LLMs are good at generating boilerplate” has always been “why are you putting up with needing to write boilerplate in the first place?”

      I really do think that it’s emblematic of a problem that existed long before LLMs: a lot of developers were/are deeply uncomfortable with stepping outside the confines of their IDE, and thus would basically never fix workflow/DX issues unless said fix came packaged as a feature by whichever megacorp was feeding them tooling. Scripting is all but unheard of outside of things like builds or CI.

      I see this divide in my own organization, actually. Our mobile developers aren’t really all that comfortable with stepping outside of what Android Studio/XCode can do and the workflows deined by Google/Apple.

      Our backend team is very different. No one uses IDEa, and we’re all Unix terminal folk for the most part. Lots of scripting and tooling we’ve created over time to streamline our development.

      So when me and few other backend engineers decided to write a small Caddy file to set up a reverse proxy to enable the mobile emulators to talk to our locally-running backend (rather than a shared non-prod backend that the mobile teams always use, in order to facilitate more efficient end-to-end testing for features), we may as well have been talking an entirely different language when describing it. The very idea of creating your own tooling to solve a problem with your job is just entirely alien to them.

    • IchNichtenLichten@lemmy.wtf
      link
      fedilink
      English
      arrow-up
      13
      ·
      24 hours ago

      It’s everywhere. Can’t code for shit? Just use a LLM bro. Can’t create art? Just throw your prompt at the latest slop machine. Too lazy to write anything? We can help with that. Women too stuck up and woke to talk to your nasty self? Guess what, our LLMs will tell your just how special you really are.

      Sometimes I just want to go live in a cave.