Cheaper to Write, More Expensive to Trust

A friend messaged me lately: Where do you see IT going? Any new patterns in the post-AI era?
My answer started with two recent jobs, both from my own desk. The first was one afternoon, when I asked an agent to build me a small UI that exports data from my workout tracker and from my Fitbit Air through the Google Health API, and merges it into a training agent I’d built. It works, and it will never need to survive me.
For the second, around the same time, I spent a week updating one outdated JavaScript library in an emergency medical dispatch system (the kind that follows an ambulance from the first call to the hospital doors, with the patient’s status reported along the way). It was already running. An agent helped me debug it and swap the library, and the rest of the week went on reading, tracing, testing, and asking what would break and who would notice first.
Same tools on both, but completely different jobs.
First, a word on what “AI” means here, because the word does a lot of dishonest work.
Ed Zitron made the point on The Diary of a CEO : the label is kept deliberately wide, so protein folding, robotics and chatbots all sit in the same bucket. Criticise the chatbot, and someone reminds you that AI is curing cancer. That same chatbot you pay a monthly subscription for has cured nothing, but it’s very happy to take the bow.
So in this post, AI means LLM agents writing and changing code. Nothing more glamorous.
Back to my friend. Somewhere in the middle of my messy reply, I’d typed: “we need to differentiate between ‘one-shot’ products and ‘stable’ ones.” Simple enough, yet most of the noise about AI in IT, the hype and the doom alike, comes from people describing one kind of software and drawing conclusions about the other.
In January I argued that AI made expertise-shaped output cheap . Steven Bartlett said it more bluntly on the same podcast : “if an AI can do it, the ‘great’ thing is not of value.”
That cheapness is real. It just doesn’t land the same way on both kinds of software. AI collapsed the cost of building software, but did very little for the cost of keeping it.
Two kinds of software
My workout UI is one-shot software: a short life and a small blast radius. So are a prototype, a demo for Thursday, a migration script that runs once, an internal tool someone needs for a quarter. Nobody is on call for any of it. When it has done its job, it gets deleted, or it rots quietly in a repo and nobody minds.
Stable software, on the other hand, is everything that has to run for years, carry money, carry people, or outlive the person who wrote it. The dispatch system that decides which ambulance goes where is one, and so are the bank core, the booking engine and the flight software.
The economics are different. For one-shot software, writing the code is most of the cost, so making the writing cheap changes everything. But for stable software, writing was never the expensive part. Changing it safely and debugging it at the worst possible moment is where the money and the nerves go.
The split itself is old. Prototype versus production has been discussed and analysed for decades, from Fred Brooks’s “plan to throw one away” (1975) to Manny Lehman’s laws of software evolution (1980). What’s new is how cheap the throwaway side has become.
The obvious objection is that the line blurs. Prototypes get promoted to production all the time, and nothing is more permanent than a temporary fix (most of us have inherited one).
That’s true, and it’s exactly why the split can’t be left to drift.
Someone has to decide, at the start, which kind of software this is… and that someone has to still be around when the prototype gets its first real user.
Fast, cheap and good enough software
AI deserves its full due here.
That UI took an afternoon. The sales pitch for AI is that building gets fast and cheap, and for this kind of software it’s true. Ideas reach a working demo while they’re still fresh, and a small team can now cover what used to need a bigger one.
It’s also the one place where I happily let the agent think for me. If its code is wrong, I find out quickly, I throw it away, and nobody’s evening is ruined.
Programmers have been here before. In 1957, plenty of assembly programmers swore they would never trust code that a compiler wrote for them. They were wrong. By 1958,
more than half the code on IBM machines
was generated by FORTRAN.
Refusing agent-written code today would be the same mistake, at least for software where a wrong answer costs you an afternoon.
Stable software, however, is a different story. The trouble starts when the same accelerator gets bolted onto it. Google’s
2025 DORA report
, a survey of nearly 5,000 professionals, found that more AI adoption went with higher delivery throughput, and with lower delivery stability.
The report puts it plainly: “AI doesn’t fix a team; it amplifies what’s already there.”
How that amplification works is boring. The agent multiplies the volume of change, and all of it flows into the same reviews and release process the team had before. If the pipeline was solid, you go faster. If it was held together by two people who know where the bodies are buried, you now break things faster.
Even a solid pipeline pays in another way. Matt Asay put it in one line in InfoWorld , quoting a commenter: “It doesn’t get easier, you just get faster.” And the time you save gets spent on more scope, which quietly becomes the new normal.
Stable software needs owners
Simon Willison has probably spent more hours with coding agents than almost anyone writing about them, and he likes them, yet in a recent note he wrote:
“The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder.
We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge.”
This is where the 1957 parallel breaks. Compilers earned trust because they’re deterministic: the same source gives the same output, every time, and you can verify it.
An LLM is non-deterministic by design. Ask twice, you get two plausible answers.
So review and verification don’t fade away the way hand-checking assembly did. For stable software, they become most of the work.
Verification also assumes someone understood the change in the first place, and that depends on how the work was handed to the agent. When you outsource, you give it work you could do yourself, so you can still check what comes back. When you offload, you hand over the thinking too, including the thinking you haven’t built yet, and you’re left with nothing to check it against.
A delivery pipeline can’t tell the two apart, because both end up as the same pull request, with the same green tests and the same velocity on the dashboard.
A recent
controlled study
put junior engineers in front of an unfamiliar library: those who used AI to understand the code scored 65% or more on comprehension, those who delegated the code entirely scored under 40%. You’d only see that gap in a quiz, or in production at 2 AM.
Nobody cares about that in one-shot software. But in stable software, offloaded code is debt, and whoever is on call pays the interest.
My week on the dispatch system was outsourcing. The agent helped a lot. Still, the week was a week, and most of it went on the part that made the change safe to ship. That part was mine.
That’s the real requirement for stable software. Every change needs someone who can explain it, defend it and fix it. Human review alone falls short of that; the 737 MAX’s MCAS1 and CrowdStrike’s 2024 outage2 were both built, checked and shipped by people. What delivers it is structure. Whoever merges a change also gets the call when it breaks, and every review asks “why” before it asks “does it pass”.
Regulators have noticed and are moving the same way. In the EU, the revised Product Liability Directive treats software, AI included, as a product from December 2026, so the organisation answers for it too.
But someone still has to run all of this. That’s what I meant when I told my friend we need real craftsmen to orchestrate the agents.
Generating code is the easy part now, so the craft has moved to directing the agent, giving it the right scope, the right context and a sensible order of work, and then to everything after: tests, rollout, and watching the thing once it’s live. The teams that master that whole process will do well with AI. The ones that only count how much code comes out will mostly produce code for someone else to own.
And when generated code does break in production, the person fixing it has to cross code, infrastructure, data, and the business rule nobody wrote down. That’s why I expect the polyvalent problem solver to come back. After all, the agent that wrote the code won’t be on call.
Who will keep stable software running in 2040?
Those problem solvers have to come from somewhere.
Look at who gets which kind of work. One-shot software is usually what lands on a junior’s desk: the internal tool, the prototype, the script. It’s also exactly where offloading pays off and nobody checks. The stable systems, though, stay with the seniors who already know them, and that’s where you learn by breaking and fixing things.
The hiring data points the same way. A 2025
Stanford study
by Brynjolfsson, Chandar and Chen found that employment for 22 to 25-year-olds in AI-exposed jobs, software development included, fell about 16% relative to the least exposed jobs, while it kept growing for those aged 35 to 49.
The timing overlaps with the post-pandemic hiring correction, so the cause isn’t settled. But the direction is hard to ignore.
None of this means juniors can’t learn with AI. In fact, the evidence says they can: in
an earlier study
, the least experienced support agents gained the most from an AI assistant, becoming around 34% more productive.
AI can be the patient tutor most of us never had, as long as there are real systems to teach on.
I know how “it works for old people like us” sounds. Every generation of programmers has said something similar about the next tool, and mostly been wrong. My worry is narrower.
You can only outsource safely what you could have done yourself, and every senior I know who outsources well learned by doing the stable work by hand, scars included. If the next generation never gets near that work, in 10 or 15 years there will be plenty of people offloading to the machine, and very few left who can outsource to it.
So here is the tidier answer to my friend. AI split our work in two.
For software built to be thrown away, like my workout UI, it’s the best tool we’ve ever had, and we should offload without guilt.
For software that people depend on, like that dispatch system, it raised the bar: more review, more verification, and more people who can debug what nobody wrote by hand.
I knew where to look during that week because I’d spent years breaking and fixing systems like it. The juniors coming up can learn the same, if someone still lets them near the systems that last.
Whether we do is a decision, and it’s ours.
MCAS, the Boeing 737 MAX’s flight-control software, could push the nose down repeatedly based on a single angle-of-attack sensor. It contributed to the crashes of Lion Air 610 (October 2018) and Ethiopian Airlines 302 (March 2019), killing 346 people. See the US House Transportation Committee’s final report (September 2020). ↩︎
On 19 July 2024, a faulty content update to CrowdStrike’s Falcon sensor crashed around 8.5 million Windows machines worldwide, by Microsoft’s estimate, grounding flights and disrupting hospitals and banks. The update passed CrowdStrike’s own validation checks. See CrowdStrike’s root cause analysis (August 2024). ↩︎