Coding Is Not Solved
Discussion 380 comments
I guess time will tell if the consumer will adapt to the lower quality of products, allowing companies to justify the existence of lazy and incompetent developers, or if the consumer will push back, forcing companies to increase the quality of their developers.
Note: I use AI every day and it is entirely possible to create high quality software with it, so long as you are not lazy and incompetent.
One reason I'm reluctant to hand over all of my work to the ai is I don't want to forget how to program or let my skills deteriorate. Another reason is I don't want to become dependent on ai and find myself in a situation where I'm not able to fly/navigate/land the airplane if my auto-pilot or ai malfunctions or fails.
Then the last reason I don't want to take the lazy approach: When I've done "one shot tests" a lot of times the ai will try and take some lazy half-ass shortcut that we would not accept if it were a human doing the work. A lot of times it just doesn't do what you ask it to do.
Where I've found ai extremely helpful though is asking questions. Asking it to build me a function that takes in a, b, c arguments and spits out x, y, z.
AI really is one of the greatest things mankind has ever produced, but I don't think it's so good yet that it can replace humans completely. Using it as a form of leverage though I think is what people should be doing. I suppose we'll see what happens to developers who let the ai take over completely. Some people are arguing that if you don't let the ai takeover completely your career is doomed, but personally I think you might be doomed if you forget how to fly the airplane by hand.
Brings to mind this classification https://en.wikipedia.org/wiki/Kurt_von_Hammerstein-Equord#Cl...
"""I distinguish four types. There are clever, hardworking, stupid, and lazy officers. Usually two characteristics are combined. Some are clever and hardworking; their place is the General Staff. The next ones are stupid and lazy; they make up 90 percent of every army and are suited to routine duties. Anyone who is both clever and lazy is qualified for the highest leadership duties, because he possesses the mental clarity and strength of nerve necessary for difficult decisions. One must beware of anyone who is both stupid and hardworking; he must not be entrusted with any responsibility because he will always only cause damage"""
Now instead of 90% stupid and lazy (harmless, useful for grunt work) you have 90% stupid and hardworking (aggressively causing damage).
You see this already, LLMs are a lot more reliable in statically typed languages with strong memory guarantees (like typescript or rust) than in weaker languages.
IMO the only way LLM code can avoid most of the pitfalls of human code is if we make new programming languages targeted at being used by LLMs exclusively. Think of languages with very strong methods for formal proofing and stuff like that.
The problem is that even if said language was invented, it would still fail catastrophically when integrated with systems not made in said language. We are very lucky that relational databases already provide a somewhat high level of formal proofing in this regard.
Said language would be impossible to parse by humans, kinda like assembly where you can parse what an isolated piece of assembly code is doing, but if you can't comprehend a somewhat large pure-assembly codebase as a whole.
It’s common advice to wire in deterministic feedback to your workflow with LLMs - static languages aren’t inherently better for LLMs, it’s that LLMs produce better code when given deterministic feedback, such as compiler results.
Based on what? This will not happen!
If I was to employ them to review the code without giving each the same baseline multi-page prompt, they go into endless loop of "improvement" with no end goal in sight.
More and more frequently, I instruct frontier models to stop and go back to the task at hand.
Luajit is under 80,000 lines of code.
Where does all that supposed productivity go?
https://innovationgraph.github.com/global-metrics/git-pushes
Just because you haven't installed new software doesn't mean that new software doesn't exist.
In fact, having more churn can lead to worse software due to diverging patterns and inconsistency
No new browser, no new iOS clone than runs on Android, no new easy to use DaVinci, no new CAD suite, no $5 SolidWorks clone, no redesigned K8s, no 10x performance speedup in Linux kernel.
It just gets “reviewed” by an LLM, which will find a nitpick while ignoring the huge fire in the core of the design, force the planner to make even more sloppy code to cover for an irrelevant test case. Rinse old tokens and repeat until you hit limits.
For example, I recently got brought in to help with quality on a large-scale system that had been ported to a new platform with the help of coding agents. The project was completed and declared operational in record time, but soon after the business discovered that:
1. The promised scalability improvements did not materialize. Instead, it got worse.
2. Observability had been lost. The telemetry was no longer trustworthy.
3. Users stopped trusting it because it was producing incorrect outputs.
What I ended up discovering was that, while it scrupulously kept existing automated tests passing, any behavior that wasn't explicitly covered by a test was free to change any which way. And there were plenty of small things that weren't explicitly covered. Perhaps because the original authors thought they were so obvious and commonsense that they didn't need one, perhaps because mistakes happen. The why doesn't matter. The point is that reality is messy and imperfect, so giving someone a chance to look at things and think, "Huh, that's funny..." is an essential part of defense in depth.
The real worst part was, this whole replatforming was a huge waste of time, anyway. The improvements they were looking for could easily have been accomplished with some controlled incremental changes to the original system. Mostly just removing a few basic and well-known performance antipatterns.
But way back at the outset, the person in charge of the project asked their agent, "What's the best way to X," and the agent gave them a trendslop answer about how Y alternative technology is more scalable and we should just port to that. It was convincing and they were under intense time pressure to just ship some code because leadership is bought into the AI hype and now has the patience of a 4 year old, so they just went with it.
A couple more step functions in model capability of the type we've seen in the past year, and there will pretty much be no reason for humans to be involved in the development process at all. All humans would need to do is communicate clearly what needs to be made and flag problems as they come up.
Kinda what i'm doing already, but for the young startup I'm at that's surprisingly tons of work. I miss the days we wrote code by hand boy those were fun 8.5 hours workdays.
Sounds like the easiest thing in the world: I wonder why did we not think of it earlier?
It didnt go wrong
And if it did, it was because you werent using the latest model.
And if you were, it was because you didnt have the appropriate guardrails.
And if you did, it's because you didnt have AGENTS.MD.
And if you did, it's because you didnt prompt it properly.
And if you did, it you're still going to be redundant soon because I'm sure the next model released will fix whatever went wrong.
I mean, sure, I could have predicted in what ways an LLM would fuck up, but there's just so many ways I can't keep up.
We just had a major production issue because someone's LLM wrote queries against dev databases. Which are very obviously dev databases because they are labelled with dev in the name, and in the table descriptions. AI reviewer didn't catch it, neither did the human reviewer for that matter.
Just don't. Fire them! AI is better than a thousand devs. What you need is testers that know what to test that AI can't, not code or UX/UI (not talking about playwright here) but business intelligence if that is testable, the things that produce results (profits) and the reason it was asked for in the first place, to solve a problem
If the problem was asked wrongly, the result will be wrong too. Fire devs, then PMs, then IT Managers if they really don't know how to outperform AI, and that's exactly the point, they won't be able to do it in code or tests or reviews, only in intelligence, for now...
> I hope you can see the stupidity here if you expect to see any deterministic results at all.
Are you expecting humans to be deterministic in the code they produce?
A human who knows 1+1=2 can still say “3” because they misread the question, misspoke, were distracted, or made some other cognitive error. Likewise, an LLM can output “3” because the generation process selected an incorrect continuation. Those are both errors in producing an answer, not evidence that 1+1 somehow has multiple answers.
So yes, human mistakes and LLM sampling are mechanistically different. If your argument is that LLMs and humans can both make mistakes, then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
It's not, I'm just pointing out that LLMs won't make that mistake.
You could ask an LLM what 1+1 is, and the number of times it says "3" is so small that it makes no sense to worry about it. It will phrase the response differently each time; that's the nondeterminism. But it won't say "3".
> then major question here is why are we building out huge amounts of infrastructure at unsustainable spending levels to enable LLMs to make the same mistakes as humans.
Yes, if we ignore everything else, that seems like a reasonable question. But let's not ignore everything else, like the fact that LLMs are much more productive than humans and likely already make fewer mistakes than the average programmer.
And?
The p(that kind of error) is pretty small now. At what point does a probability coming out of an LLM look like "knowing", such that spitting out the wrong answer despite that probability looks like a health problem, a typo, or even just boredom? (Thinking of the Lizardman constant here: https://en.wiktionary.org/wiki/Lizardman%27s_Constant)
It's a continuum for both them and us, even if the mechanism is wildly different.
> Making mistakes is not the same as non-deterministic.
i.e. when the dismissal is "non-deterministic" when it should be "Making mistakes", is itself a mistake.
But this is plainly false. This kind of unforced error occurs all the time.
For example, once when I was in high school I traced an error in my math homework to an intermediate calculation of "2 + 2" as being "3". There was no reason.
What we can say about humans is that, if they know that 1 + 1 = 2, (a) they are unlikely to change their mind about this in any kind of lasting or permanent way, and (b) the rate at which they will mistakenly produce other values for 1 + 1 is very low. But it will happen occasionally, and when it does happen, "they just suddenly decided on the wrong value" is an extremely accurate description of what that looks like.
The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.
As humans we don't have our memory reset multiple times per day
I've seen humans vote for Brexit, re-elect Trump, ask questions clearly already answered in an FAQ, try to pull on a door labelled "push", and insist on giving me homeopathic silicon dioxide pills* that cost £5** for a 10-12 gram packet.
Continual learning is a difference, but not by itself a reason to care about "deterministic results".
Nor, indeed, correct results.
> As humans we don't have our memory reset multiple times per day
Humans need sleep well before they can read a million tokens' worth of written text. We're more like 300k tokens if you're actually reading and not skimming for 16 hours straight.
Again, different (in soooo many ways), but this isn't a relevant difference when the topic is "deterministic results".
* yes, sand: https://dailymed.nlm.nih.gov/dailymed/fda/fdaDrugXsl.cfm?set...
** and that was what it cost in the 90s
There is no RNG involved when I decide to push vs pull the unlabeled door to my building every morning, it becomes deterministic because its baked into memory
You can put stuff in context to deal with this but you can't do that for everything, its not practical and you would blow the context window
How does the system behave in a variety of scenarios including failures and restarts. How is state maintained coherently. There are the kinds of systems problems that an engineer needs to reason through, and if there are bugs in such decisions, they end up becoming costly. I dont expect AI or LLMs to solve these problems at all, since each of them has nuances and tradeoffs which are specific to each system. In short, there is specification complexity in precisely describing system wide behaviors, and unfortunately, there is no lean/tla+ to meaningfully describe systems at scale. You could then ask: How can a system have guaranteed behaviors if they cannot be even stated or proved formally ? The answer to this is how protocols like raft/paxos initially convinced us of their behaviors which is in human review and understanding. That begs the question: How can human review and understanding be reliable, and the answer is that it is not reliable, but humans have ability and processes to continuously learn from experience in the real world. So, our understanding is grounded not only by whats out there in books etc, but also by our own interactions with the world.
Long story short: The responsibility for system-wide behaviors of software systems relies on human review and understanding, which while imperfect can continuously learn.
Yeah. To me it seems very much like the "use dynamic typing for everything" fad. You had a bunch of junior and/or incompetent developers who went around insisting that type declarations are bad, static typing slows down development, you just code so much faster if everything is dynamically typed. And in the context of a new project, they were totally right. It took a few years for the debt to finally catch up, and people realized that these massive, untyped monoliths they had were unmaintainable. Now the two biggest dynamic languages (Python/JavaScript) are effectively typed languages, because nobody uses their untyped variants for serious work.
Dynamic typing still has great uses -- interactive data exploration, putting together quick scripts (though less relevant with AI...), or even just simple prototypes -- but what we tried to do with it at the start, as an industry, was clearly dumb as hell. I suspect we'll look back in 5-10 years and realize that with some of the stuff we're doing with AI, too. It's already happened with things like Gastown.
I totally understand where this is coming from. I too am struggling with accepting that my 30+ years of programming experience is quickly becoming obsolete. I'm losing sleep about this, it's tough.
But just go ahead and give the latest models (Opus 5.5 / Astra 6 as of today) another try. See what they are capable of and read the code which they produce. Any problem area, low level C++ or high level Typescript or Clojure or a weird combination of these..
Don't be shy, give them a big task, let them build an entire app, UI and all..
Now compare the output to Opus 4 or gpt-5 from 1 year ago - when they couldn't put together a single function without it being weird and buggy.
This is exactly my problem, not that the models are very good already, but how fast they got so good. So if coding is not solved yet, it'll get there very soon.
I don’t see the ops comment as a rebuttal. He agrees coding is not completely solved. However, it’s getting closer to being solved.
What kind of coding are you using these models for? Most of the people I know who share your perspective never go beyond the prototyping stage. I’d be curious to hear from anyone who’s been AI coding for more than six months, shipping it to real users, and isn’t looking at their code at all.
Feels like I read comment similar to this one each year since 2023.
That's when the models started to be coherent enough for real work.
They still fuck up, but it does not feel the code was written by drunk interns anymore.
This reflects my own experience with these models. It's not a matter of inflection point if you ask me, it's a matter of accruing capabilities last year the output was not up to my standards 98% of the time, now it looks more like 30% of the time.
I am sure next year models will be better, but the point where the models begun being good enough to start using seriously for my use cases has now passed.
I have found that a willingness to look like a temporary dumbass (primarily to yourself) is the largest predictor of success with pretty much everything.
What are the consequences of asking an LLM for the moon and receiving low earth orbit instead? Who cares if the proverbial rocket explodes on the pad? This is all happening entirely in a computer system completely under your control and likely at relatively low cost. No one else has to find out about your mistakes if you don't want them to.
Now the syntax is handled for you, you have a research assistant, and someone that can really dig through the details for you.
The rest ... is still there.
No, I really am not. I'm doing maybe 10% of what I did before. The rest is filled up by other, usually higher order tasks like planning, product management and work orchestration.
I feel as though we are all doing the work of 2 people + an Architect, not so much 'Product' issues, although I'm sure it varies.
But I can't imagine how any actual software is written with 10% of the effort.
I’ve been using AI as a great pair programming partner for about a year now, but every time I prompt it to write code agentically, it just makes a pile of crap that I end up spending more time fixing. This is so very quickly becoming not the case anymore.
I think core engineering skills are never going away (I’ve been doing this for 25 years and come from a background doing C++ for video games), that you will always need to have a mental model of the code and if you’re going to call yourself “professional” you need to be able to go in and fix/build by hand. But you’re going to get left behind if you aren’t at least willing to eat some humble pie and re-evaluate your views on agentic coding every few months as these tools get better.
Grieve, I know I have, but there is joy on the other side. I used to love getting lost in flow state with Soma.fm playing and the phone unplugged, that is still there, but what that looks like is changing fast and table stakes in this industry has been “adapt or die” for as long as I’ve been in it.
The one exception is for stuff that I care very deeply about. For a very small subset of projects where I'm willing to spend hundreds of hours ensuring that what I build is the best possible thing, AI still hasn't been able to match my work unless I micromanage it, but at that point just writing the code myself is actually faster.
For almost all jobs though AI is probably the future, most software developers never cared that much about their company's code anyway.
Well, for one, they're capable of draining our (or companies') wallets.
I resolved a huge merge conflict for $60 today. Opus 5.5 did a great job and spent just 1h 16min on this. I could probably run six such sessions today, if I disregarded the need to read and understand the code.
This money has to come from somewhere and my concern is that it will be from decreasing the number of people hired and/or their salaries.
At the same time I firmly believe people who had a tendency to produce tech debt will keep doing that, regardless how brilliant LLMs will become. Unscrewing this is going to cost a lot of money.
i guess another way of saying this is that on the micro level a lot of this stuff is not rocket science, but at the macro level it becomes hugely complex.
This is not a good premise. All over law, you will find people made responsible for what they don't control and they kind of own. Unleash a dog that harms a child, or just have it in an environment where it can escape, and see what happens.
There is such things as unpredictable situations where one might not be held responsible, as a problem might occur well past reasonable guidelines.
So of course you can be held accountable for what an AI that uou supposedly cannot quite control does, or for the AI-written code you deliver. Treat it like the releasing a wolf pack, or selling an unsafe toy that can maim children. There's precedent everywhere.
The difference seems to be that some companies are above the law apparently.
Can anyone tell me why we have 40 or more programming languages, with about 10 popular ones? Then about 20 frameworks in each of them. And add another 200 popular libraries for each language? This matrix make no sense till you realize - it is preferences all the way down.
Most of us engineers have built our own mental model of programming. We are all right. But the users do not care. LLMs are here to produce code closer and closer to the metal as needed. They can sit and create a graph out of every spec, use an AST that they develop and run on the CPU if they have to. They will do it. No amount of us discussing will stop that.
Programming is going to be re-invented. I do not think the current ways to write software will even matter.
Because too many of the rest of you lack the good taste to use Lisp.
No you're falling victim to the common programmer fallacy that "my use case is everyone's use case"
These things exist because people had different use cases and priorities over the years
... Oh by the way, did you know that Claude Code is actually a mini game engine? /s
Those who claim LLM-generated software is good enough:
Haven’t written code in ages
Cannot spot if their code figuratively had 6 fingers!
Have a low bar for what good looks like
Don’t care about quality or NFR
Have difficulty understanding an S-curve
There are exceptions like antirez, but I think this does hold for many loud optimists out there.Costs which largely go away if the software you're using is so custom that it's only applicable to your six person team anyhow. That wasn't feasible before, but it is now.
The previously fashionable one-size-fits-millions approach to software hasn't treated users well enough to expect them not to defect when they're suddenly able to go it alone. For many, the quality issues are worth tolerating, because they're still less painful than something which was designed to be sold rather than to be used.
My issue with it, is that it gives you a "lazy" option every time that doesn't require the same level of thinking. I understand that this is completely on me as the developer, and the simple solution is that I need to make sure I'm taking my time to learn and understand what exactly the LLM is producing. I try this and have set up separate skills to make sure I'm building my understanding as I go.
Regardless, if I sit down today and implement something without the use of LLM, it takes me a lot longer, but once I get into it, I find a state of flow that I can never get from the back and forth reading of LLM output. Then when I finish, even if my solution is not perfect, I have learned so much more and my own context of problem is so much better, where usually then I can review with an LLM. This usually leaves me with a better implementation and more importantly one I can stand over. I think for a newer dev like me (~2 years experience), since I haven't built up years and years of problem solving experience, if I don't carve out time in my day to put down the AI tools and improve on my problem solving, I'll plateau and that's my biggest push against all this LLM use. I don't necessarily disagree that 'coding is solved', to be honest, I think it largely is, but it's still the foundation for me to be a good Software Engineer and I definitely haven't solved it.
OTOH , a lot of people, including people that should know better (looking at you, Netflix) are definitely holding it wrong.
In my experience, managing LLMs is much easier than managing an office of junior engineers, And more productive at 1/10 the cost. Now where the next crop of wise seniors engineers is going to come from, well, that’s a different problem.
LLMs are like 7 year olds with PHDs and coke.
I develop mission critical firmware using AI agents. I run a full agentic office, 10-50 agents at a time most days. If you don’t have almost as much documentation as you do code, you’re probably going to have a bad day. Documentation driven development is the happy path.
You need docs on your coding standards, your review methods, your test coverage standards, your protocol specifications, your build plans, the plan delta/decision matrix, user stories, etc etc etc.
In the LLM age, documentation is code at the highest level of abstraction. The LLM is a transpiler.
The loop is constantly planning, naively reviewing of the plan, implementing, test coverage, contract review, naive review of the delta, delta of the delta fix, test coverage review, maybe repeat back a few steps, then finally passing the proposed fix off to engineering governance, which maintains the standards docs. Engineering review doesn’t write code, it gates merges. It might be accepted with fixes, or it might be rejected as worse than the problem it solves.
The key concept here is that the documentation and the code are reviewed together. Where they are non coherent one or the other must be resolved. That’s where human judgment steps in when engineering governance isn’t absolutely sure of the intent. The docs are always kept coherent with the code.
The “one weird trick” that makes it work is -incentive management-. It costs nothing for gov to send work back to the drawing board. The agent’s try really hard to avoid that outcome. If gov had to write the fix , half the crap would pass right under the radar.
This works because I emphasise compaction as a metamorphosis that involves a loss of continuity for the model, “old you, new you” and they go through an abbreviated version of the onboarding after compaction (a good idea anyway) so they budget tokens like it’s the elixir of life. It’s weird but it works. Also, totes worth it to ride through the compaction. A totally fresh model can take half a day to really be in the groove.
- Coding in the small is solved. I have a current state, I want to change it, and I know how I want to change it. Eg, I have a blocking TCP handler for some reason, and I want to make it async. I can either fiddle with it or just let LLM make the changes for me.
- Coding in the larger sense is never solved. You need judgement to decide what you want made. No matter what you're building, there will be decisions to make (Who/what is it for?) and those decisions change over time. LLMs can take some default decisions for you, and if you're fine with those, you get the default (great for POCs). However you might not even realize what it decided to do for you. At some scale, you will be spending a lot of time going over those decisions. But what we have now is that the friction of changing the decisions is quite a lot lower. You can now test a lot of things that previously were very time consuming.
- The point that LLMs are probabilistic is not as important as it's made out to be. If I ask a junior dev to code up something, I also don't know what he'll make. Heck, you can be sure that you are able to solve something, yet you yourself don't know what the solution will look like. Maybe it turns out the library you were going to use isn't appropriate after all. You don't know what you will use in the end, but you do know that something will fix the issue. There can be more than one solution to a problem, and it doesn't always matter which one you find.
- I STILL think that LLMs are at their best mostly as advanced predictive text. In the sense that it's mostly good at implementing things that you've decided are needed. This can mean a heck of a lot of code, but you have to know the tradeoffs. What was decided, what were the costs of those decisions in terms of maintainability, money, time to change it, and so on.
I’ve been through quite a lot of programming eras over the years (compiled, interpreted, loosely typed and finally JS/Python for everything), but this time I just can’t adapt anymore.
I’ve jumped this sinking ship nearly two years ago and while I was skeptical about it at first, I’m now more and more happy about this choice.
Just a simple reactor, my laptop's only little. But still.
/s
It isn't solved because they cannot, in fact, do what you claim. LLMs write code worse than humans do, even "frontier" models.
I have no doubt that if you provide any AI system with an oracle with expected behavior that it can match that oracle with some amount of $ and tokens. I haven't seen any demonstration of anything else. Rewriting a codebase was always a challenge for humans not because of complexity, but because of the time and effort involved in matching the old version's prior behavior. It doesn't have anything to do with the serious level of work required to build something truly new from scratch in a performant way.
Make it port some Rust to JS, see how that goes.
And Rust isn't a super difficult language to use on a daily basis. Sure, the initial learning curve is obnoxiously steep, but it's arguably easier to use than other languages once you get past that.
The LLM crowd have this habit of equating something that they don't understand with requiring some kind of advanced skill.
https://georgzoeller.com/blog/posts/what-reverse-engineering...
For example, any amount of software development involves fixing bugs, getting feedback from users on ideal workflows, an iteration loop of performance and bug tuning, etc. AI cannot simply create, from scratch, perfect software. Even using the SOTA models on max effort does not produce bug free software of any meaningful complexity or innovation out of the box. All that has changed is that the act of physically writing code and implementing existing patterns is now effectively a marginal cost.
Most line of business software is not e.g., delivering a company's income. Most software is in back-of-the-house internal products that do various internal tasks. I have no doubt that these processes are now far easier to build.
If the new Copilot is so great, why is it completely out of the current zeitgeist when compared to Codex and Claude Code?
GitHub's Copilot cloud agent offering is suffering with a case of some of the worst corporate ADHD I've seen. We built a cloud agentic development pipeline on it, and it seems like almost every other week they silently change something with zero public announcement or documentation that creates real disruption for our team.
That's real, breaking changes to the platform that clearly aren't being tested/reviewed before being pushed to prod. Again with zero public announcement or documentation.
Support is useless – we're paying customers in the 4-5 figures and our tickets go unanswered.
Especially expensive when you take into account the amount of that code which must have been boilerplate & meta-code in nature, meaning it should have been straightforward to move.
The most shocking thing I've read in this entire thread. Are SWE salaries really that low across the pond? That's entry level in the US. So 90k GBP a year all in, including benefits, employer-paid taxes, seat licenses, hardware, furniture, HR? So what's the after-tax takehome for someone senior enough to convert an enterprise codebase to a new language over one year?
When the Go team ported the original compiler from C to Go, they wrote a program that did ~99% of the work
"write a program to transpile A to B" instead of "port A to B"
would be curious to know how many times "unsafe" appears in there, have seen rust devs comment on how the ais like to use unsafe to work around difficulties with memory management, like how they will sometimes subvert tests
But 'coding' per sey is 100% solved by LLMs - they write compiler perfect code all the time.
The question is not 'what it writes'.
The LLM is like a writer's assistant, who has perfect prose and grammar, but doesn't really write 'stories'.
""The reason LLMs are successful in writing code is because we’ve made a feedback loop that feeds the syntax/runtime errors back to the LLM and loops until most errors are solved or hidden."""
No - LLMs are 'good at code' because they have been ultimately 'trained' by the compiler.
All of the various SFT/RLHF methods etc. are using the compiler as the verifier.
I wouldn't say this at all. They mess up regularly. What makes them able to write syntax that compiles is their harness which verifies the syntax and provides the feedback loop for them. But LLMs will happily give you code that doesn't compile.
I used to write code by hand, literally hand, I used vanilla vim with only syntax highlight and line number enabled, no more. So I know I can write code. I wrote C and Python most.
I used to vibe code, I am still vibe coding, multiple projects at same time. I use opus, astra, luna, deepseek v4.1 flash, I use claude code, codex, pi. I vibed a project to manage my coding agents' sub agents, skills, agents.md file. I definitely know how to vibe code.
The vibe coded project seem working fine.
(Declaration first: I don't advocate cryptocurrencies, I never liked them)
Until this week I vibed a crypto wallet, with many open source wallet code available for llm to train and learn, astra designed the project and wrote spec.md, luna implemented the project, astra and opus then reviewed and fixed issues, multiple rounds.
I feel confident. I import my private key.
I interact with a web3 app.
Error. A bug astra and opus missed. I didn't read the code, I vibed it, I don't know whether my money is lost.
At that time, when your money is at risk, you know you should have read the code yourself, you should know what happened to your money instead of asking a lllm to debug it for you.
All this doesn't change the fact that software engineers are going nowhere because nobody trusts AI. If a model can escape highly secured sandboxes, then we're definitely not running these agents overnight on our systems. I am sure the next-gen of models will focus more on security and the trust factor will start developing, but that's a long way down the road.
People trust people, not systems.
As for accountability, it always laid with the employer. You think those nameless contractors whom Boeing hired suffered any consequences for that 737 Max glitch? Using AI won't change that.
AI doesn't have to solve all these coding problems to be worth handing the reins to it: it just has to substantially better on average than humans over the long haul, which it already is, especially if you have good verification of "done" and "working" in place through automated testing mechanisms. Perhaps we might say that QA is having its moment.
It doesn't mean humans aren't needed, but they aren't writing much if any code anymore.
And this unaccountability is transitive. The amount of time I've seen people successfully justify issues based on the fact that Claude/Astra/Codex wrote it is absurd. And it comes from the top.
We've had AI ship made up data to clients and tech leadership was like, "haha, that's AI for you."
> If you’re toying around, LLMs do a great job.
This too. We have a lot of business guys that develop tools with Claude that look like they work, then tech team gets pressure to deploy them immediately, because they assume that everything must be a prompt away. It's not (though, we do have some wizards on the team who make this true enough).
> AI overdose is a thing and it directly puts an expiration date on your skill set.
Agreed. But I'm not in a position to push back. It's not just managers, but tech leadership who are all in that on the fact that humans shouldn't program anymore. I still fully believe that I'm better than Claude / Astra in my specific domain, but people give me a hard time when my PRs contain what look to be human-generated code.
Overall, I agree with the article, but I'm still pretty sure I'm falling behind in my apprehension to fully trusting Claude/Astra and not giving into the approach of burning millions of tokens daily to generate PRs so large they break github (two of which I approved today).
So are humans. Every codebase I've worked on has duplicated code that has been written in slightly different (but hopefully equivalent) ways, often by the same person.
I think a few of the industries listed like defense and aviation have low risk tolerance. However, from my (somewhat brief) experience of working in two health techs for a couple of years, I strongly disagree that healthcare has low risk tolerance for tech. Granted, they make run-of-the-mill CRMs, but I was baffled at how tolerable it is to have egregious user experience that makes users waste multiple hours per month with clerical work that is very painful because the UIs are very slow and buggy.
It means risk that the software stops working after an update. Which usually trades off iteration speed and best practices (i'm pretty sure the average startup has way better security practices by just delegating to google/aws than the average manufacturing software business) in exchange for a rigorous testing and rollout schedule.
So I'm also not sure that the article has a point at all, the human writing the code was never relevant to avoiding the "risk" in these industries in the first place.
I was just at the Explore DDD conference in Denver and a portion of Friday was sitting at the cafe tables informally discussing the impact of GenAI on software engineering with notable people. Most of these people were deeply concerned that if we lean into using GenAI for “everything” that our collective knowledge will dissipate. I was the vocal contrarian. There are many historical examples of humans obfuscating knowledge to simplify progress. Does anyone solder their own microchips at scale anymore? No. We have highly sophisticated robots and machinery to do that work with extraordinary outcomes. In software engineering, if you remove “coding” as a discipline you’re left with all the other aspects of designing software which I contend can be retargeted in college CS curriculum. The leap isn’t about code reviews. It’s about design reviews and that’s where better outcomes are served regardless of whether GenAI is involved or not. I have a roughly year old codebase at https://github.com/ChicagoDave/sharpee/ that is designed by me, but generated by Claude Code with my own skills and agents as guardrails. I’m fairly certain the code I extract from Claude doesn’t require human review, but the design of the system and its changes are continually reviewed by me. My contention is that we “collectively” are still trying to discern where the AI/human line is and most are still “holding” that line to human interactions. Let it go. Define what part you do need human decisions on and focus on those things.
Is this true? I’ve worked in various software companies for over 20 years now and I’ve never had to worry about risk in particular, and the codebases were all somewhat bad in areas and buggy (as tends to happen when a codebase gets large and old).
The past 6 months LLMs have crossed into barely needing to check the output territory for what I work on, and generally write very similar code to what I was going to write. It needs some common sense to use it correctly of course, like going feature by feature and keeping the commits fairly small, but my job is easily 10x less work/time for the same results.
I’d be genuinely surprised if even 1% of software requires aviation levels of risk tolerance and testing, but maybe I’m mistaken and most people are working on much more mission critical things than I am?
Instead, I think what's closer to solved and what we're in the process of solving is product development.
Story: A while ago, I had a few programmers who were really, really fast almost always missed the mark on the assignment wrong. I loved having them on projects because in the time my senior precise engineers could deliver a MVP, the fast engineers would build the wrong thing, collect feedback, reiterate, build the wrong thing, collect feedback, eventually inching closer and closer to a product people would pay for, and it would almost always get delivered faster than my seniors.
I feel AI does the same thing.
I got lazy around claude fable and astra, and asked them to work in loop (pick specified issue, develop it, qa it ...) have a separate CTO checking on arch.
at the end both models swore that the code is perfect and well designed and nothing is lacking.
I ran the software and it suddenly started writing large amount of data to CSV files instead of the typical DB usage.
AI decided to use csv for testing, and just drifted away. 0 regards to the actual project, 0 regards to common sense.
anecdotal but really weird, the project category is rather standard, I wouldn't accept such a mistake from a junior developer.
It compiled and ran just fine. If you weren’t reviewing the code holistically or keeping tight book keeping of your allocations you would not have noticed. Every single commit in isolation looks perfect. Very eye-opening
Isn't it the opposite? How to build something is rather solved, but what to build isn't?
But that's not solved in traditional product development either.
Product development an iterative process to get a product fully functional. In 2021, if you ask me what the timeline for a small product/substantial feature, I'd say a few weeks to a month to get a basic MVP, and then another 12 to 18 months to get a feature polished and in a good shape to be stable.
When people put it in the coding frame, what they do it as is saying we've gone from 18 months to minutes or days. That's just not true.
We have gone from eighteen months to depending on the complexity, a 1-4 months.
aside: To be candid though, the compressed time also means the frustrations people experience with a product in 18 months have also been compressed. They still exist, they're all there, they're now just non-stop.
So transportation is not solved either? In that case, beam me up Scotty, I can't see any hoverboards around.
"Software architecture is solved, there is no longer any need for human intervention in the architecture or design of any software, nor in any subsequent stage"
Well, that is what Dario et al would like you to believe. However, the argument they can actually defend is:
"Programmers no longer need to type the code out by hand in most cases, they can get the result they want with (several iterations of) higher level instruction."
What a farce. Too bad people eat it up!
Now SWE job is to make sure to combine and instruct the AI tools to produce ever more complex outputs, judge the tradeoffs in these outputs, and guide the tools further.
In a way, it's not that different from pre-AI coding.
Coding is not solved because you can’t simply prompt an LLM to make an AAA game or enterprise tool.
Dear lord. Is that supposed to reflect the average thoughts and motivation of a person you want to hire? Or that of their employer?
Nobody has to be in fear, but we do have an ingrained knowledge that there are consequences, good and bad, for our actions
That concept might work a lot of the time but you will definitely run into situations where that'll never produce a correct or working response. To actually learn something you need an environment/playground to apply what you think you know and observe the results. Without that you're not really learning, you're jus regurgitating what people want to hear.
AI can write CRUD API endpoints almost perfectly now. It can also write quicksort, a heap, whatever much quicker than I can.
It really sucks at designing types and apis though and when it creates types and apis it doesn't think or plan for the future way the system will evolve (even if it's known up front how the system will evolve).
I suspect this will remain a problem for the models for a long time. All the things that the models are currently good at are the low hanging fruit of reinforcement learning for coding.
Think about the kind of reinforcement learning environment that needs to be created to train a model to become good at building and designing large scale software end to end. It would be a slog because you need to build the large scale software up front and then break it down to train the model to construct it in a systematic manner that allows for the software to evolve. And then you need enough of these training environments for it to generalize. I think they will eventually figure it out though but it may take a while.
Does that really matter? Those are things so that humans can better understand and extend a code base. That mattered when writing code was expensive and took time.
Now if it can pass all the tests it’s fine. If there’s an issue just have it rewrite things immediately. New bug? Generate a new test and rewrite code.
All, or many, of the old things that mattered just sort of don’t anymore.
Much of the code needs to be performant and the LLM knows this and grinds on it. Less and less human inspection is needed.
I think 90% of software can be written like this today.
You’re simply testing outputs. Make a spec but ultimately ungodly amounts of tests can be built quickly to ensure the program is outputting the right things.
It's funny because, I think the author is spot on in this article, including the problems identified in my quote above, but none of those are the reason I personally dislike the rise of LLM code - for me it all comes back to licensing and attribution. Everything else is just a cherry on top of the diarrhea sundae. I certainly don't want unmaintainable, unreliable, insecure code, but even some of the best software occasionally falls victim to these traits (its simply intrinsic to LLM code).
What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.
If we are committed to it, if we put the resources towards it and dedicate the time to it (and we could do this just by saying: it will take half as long as it used to take!), software built by llms in healthcare, finance, automotive, defense, power plans, aviation, manufacturing can all be made MORE reliable and better with LLMs... without ever reading a single line of code. The LLMS are very good at logic, by the way.
Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while. I felt the same way in 2025. I've written 100s of thousands of lines of difficult code. You, the person reading this, has probably interacted with software I've written. For a time you would've interacted with it every time you made a debit card transaction in the united states, for example. I understand code, and care about quality, and that's why I'm all in on LLMs for code.
Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.
“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.
Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead).
But it does give you surprisingly stasble rube-goldberg machines.
And thats basically what 95-99% of enterprises want from their software.
It annoyed me to no end when i began my career, but at this point ive accepted it and can definitely still have fun developing software with llms. As a matter of fact, as my perfectionism approach to software in my earlier years was never really appreciated... So i dont really mind the new MO.
I still occasionally hand write though, esp. at the dayjob where ive got super small token budgets while continuously being told to use more AI. But that's normal, employers usually give off bipolar vibes with multiple stakeholders wanting to advance each of their bonus package KPI of any given quarter
Everything except what matters most: human time.
We're not writing theorems, dude.
Except in the equally pedantic sense that every program is a proof to a theorem...
We're writing plain enterprise and web software, closer to CRUD than NASA.
If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.
Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years.
Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.
Okay but how does AI change any of that? You can still do that with AI.
> as long as it’s sufficiently documented.
AI definitely helps with that.
> What is essential is that for every part someone did reason through it with the necessary rigor at some point.
Why is that essential though? What if the person who reasoned about it dies or leaves? Moreover, why is it imperative the reasoning happens at the source code level?
> Okay but how does AI change any of that? You can still do that with AI
With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc.
With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.
Reading the code may not be enough to understand the behaviour of your program, but believing you can understand the behaviour of a program without at least reading the high level code is truly silly.
(by high level, I mean the code living in the higher layers - of course we don't often read the code of the generated assembly, or the interpreter, or the browser, but that's because they're reliable abstractions, unlike prompts!)
have you ever used a library after only reading the README and documentation, or do you always pull the source and read through it before you think you understand it?
This is different from building a product, which you only interact with via UI buttons/CLI/etc, without reading any code to understand how it conceptualizes that product's problem domain area.
People do that latter thing, and we call them "users", not "developers".
Yes, libraries aren't bug-free, but they give me a reliable abstraction tested in the field. Not rarely you dig into library code if you notice unexpected behavior.
If we could rely on our LLM or colleague written code, or own code, have run through the same amount of requests, sure I wouldn't need to review it, as my confidence can be north of 99.9999% it works correctly. But we can't.
This sounds like a typical testimonial whose mind has become captive to Claude. It is like Scientology.
How can you not see the progress?!
LLMs are a great thing for bug fixing. However they are not a miracle. You still need to do all the other things about finding, testing and fixing bugs.
You also need to care about bugs - vibe coding rarely cares about bugs.
if you're an MBA-brained exec who doesn't actively use LLMs to code and you just believe whatever slop it outputs at first without checking it, you're not going to realize how recklessly it can be used, how you need to be critical and skeptical of its outputs, that you need to explore it's reasoning and logic (which is still really easy compared to understanding legacy code and barely takes any time!)
say you also believe all this marketing hype about 'how dangerous (ie capable) AI agents are.' LLMs can do anything you think so you just say 'ship it' without building out the tooling and capabilities to enable faster code review and better tests. and to keep the shareholders happy, you start cutting jobs that you can't directly connect to a KPI (ie the platform/SRE team who would be the ones who can trial, onboard, and maintain those capabilities for your teams)
and from this, suddenly a lot of debit card stops working and the only one getting the blame are individual SWEs trying to hit their sprint velocity. the fact that you fucked up the whole SDLC real bad with your incompetence gets you a golden parachute and you job hop to a better paycheck. rinse and repeat
I've built payment rails. Six nines SLA, high capacity, resilient distributed systems.
I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better.
Rather than debating if these models are good (they are), we should be trying to figure out if most of us will still be around in three years. You don't need a two pizza team anymore.
"Look to the person to your left and to your right. Only one of you will remain by graduation" kind of energy. I'm not sure all of us is going to be in this career much longer. We'll have to see what the demand side looks like.
on-call still exists. have fun round robin'ing that with 3 engineers.
On the other, my local pool company is hiring a software engineer and hardware engineer because with AI, they can replace a 2 pizza team as you so succinctly put it. So no two pizza teams but that doesn't mean all the pizzas are gone, they're maybe going to be spread out and not concentrated in CA, between orgs you might not have thought as "tech" before.
/s
You should seek professional help about your emotional issues.
It must have been a huge shock when you were suddenly transported from a working parallel universe into ours back in 2024.
Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.
TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.
At my current place we not only have automated tests, static analysis and static rector (linting, but also automatic pattern matcher for problematic code) but also: - architecture tests that define relationships between application layers - ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated
I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should do is kept within their heads. If we define this knowledge in writing LLMs can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (gherkin)
Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works.
You're not going to get people to stop doing that by arguing on the internet, but in the end it won't matter, because it will stop, naturally.
In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness.
I expect the exactly opposite to happen. These are going to be the most requested developers as the last ones that understand how it work.
They would then be convinced to use AI for speed, but vibecoders that just prompt AI are the ones that won't find jobs.
At least some places are abolishing formal QA because LLMs. There's a cult of speed uber alles that has a big intersection with LLM enthusiasm.
That cult was well established prior to LLMs
If coding were solved, then this would be true no?
I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?
So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).
Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?
> I think a codebase generated by AI is actually more understandable than one generated by humans at this point,
From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage, this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.
I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.
Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.
Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.
One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;
From "Groundtruth – checks your AI coding agent's claims against the Git diff" https://news.ycombinator.com/item?id=48838209 :
> "Follow up to verify that the work was actually satisfactorily completed"
> Are there other sound management practices that aren't yet effectively implemented in current gen agents?
Oh, and always write tests, docs, commit messages, and changelog entries; but don't waste tokens on documenting something that doesn't verifiably pass sufficient tests.
It's not just a false dichotomy, it's intellectual dishonesty. It wasn't that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, "just ship it!" cargo culters, or the product of management breathing down the necks of otherwise good developers, there's plenty of "human slop" running in production across servers worldwide.
Correctness has never been a priority across an industry where rapid iteration and feature delivery drive sales. There's always some opportunity cost to doing things right, at the price of technical debt down the road. If AI is primarily used to produce fragile code, people will be wary of AI solutions. There's also ongoing public debate about AI safety and alignment. Deploying AI in safety critical applications feels riskier than ever in the current environment, even though it doesn't have to be.
It's wild to read this stuff and then also deal with the constant headaches of day to day hallucinations when interacting with Claude et al.
What kind of domain are you working in?
It's a bit better than 18 months ago but it's hard to say by how much. It just seems like the culture has moved to building up fixtures that let the LLMs brute force the problems. To my eyes that's the opposite of solving things logically. It has the added effect of hiding how the sausage is made, though.
I mean, how can they possibly say they haven't written a line of code if they're actually going through it? I can only assume they're just looking at the results. So then how can they judge it's good at logic?
If it was so good at not making mistakes, why even have tests? It's nonsensical on its face.