The problem is not AI code, but not knowing about system architecture or intent
Discussion 203 comments
It boggles me we completely forgot that the world operated like this just 4 years ago
Unfortunately I learned that not everybody thinks this way. Some orgs do imperfectly fine without good engineering discipline, and that has been the case before AI... AI has only made it easier to give the appearance of good engineering, which is exactly the pre-AI goal of many orgs
And I don't think it was ever necessary to go to the point where people just gave up authorship. These were choices made by adopting the "I'll do everything for you" agentic "harness" model that shipped with Claude Code but it was never inevitable.
e.g. we completely dropped fill-in-the-middle completion OG CoPilot auto completion model. That combined AI authorship with a human always in the mix and I actually really enjoyed it. It's just that the models involved were pretty stupid. We totally could have had IDE / shell / tooling integration that kept people in the driver's seat while automating parts of the drudgery away. Instead what we got was a simple chat loop with "oh, whatever, you go do it" being the ultimate result. Cuz, you'll totally review everything after and understand it, right?
The things should end by quizzing you on what was just made and if you don't pass, just throw it away. That'd be funny to watch.
I'd think that's what they call paradigm shift, and this probably repeated across generations from the introduction of the printing press, PC, the wheel, the internet to stochastic parrots that reduced what we still stubbornly insist require our special neurons to mere statistical modelling that can be aggressively scaled.
The issue of course is that if you do invest the time, then you're no longer saving time by using AI. You're just spending it reading and trying to understand something you didn't write. And that can be unpleasant in its own way.
My hot take is that for parts of a system that can be considered its core, forming a deep understanding is almost always important, and so is knowing how the different business domains integrate and where the connection points are. For many others, a high level understanding is sufficient. The difference is that now, with AI, you can make that choice. Before, you had to write everything yourself, and for any sufficiently complex and long-lived system it became impossible to hold all of it in your head.
And having AI code to review is no different than any other code that ever was to review, so the review tooling is - as it necessitates - also benefited by lugubrious application of .. more AI. But: all AI is human reviewed.
So it's not a big impact. We just don't ship code that isn't 100% human reviewed, If that's insurmountable: you're doing it wrong. Use AI to make code readable again.
And then, also, put AI back in its box. Don't give devs 100% full-time API access to subscriptions: give them actual hardware to use, to go 100% local.
Local AI is, thus, the best AI, folks. Don't use more than you can run locally, is a great way to keep AI code properly maintainable.
The industry will prove this, itself, sooner or later: If you can't put your AI in its box for safe-keeping, you're doing it wrong, anyway... and should've already learned this practice as a habit, decades ago, vis a vis future-proof tooling... (See also: not logging everything you do with an AI? Big fail.)
Sure, the absolutely intoxicating addiction of Big Metal AI™ is going to put a lot of consumers in a deep, deep pit of Neo-Illiteracy - however: 'good' AI code is actually just good code.
> And having AI code to review is no different than any other code that ever was to review
If your company is sticking to "everything needs (human) review" then you wont run into most of this. This issue is that a lot of companies are using AI as an excuse to remove that review process (either partially or entirely).
Everyone can write code these days. Trouble is, all code has to be SIL-4 code now, because, human, you will never know if your compiler trusts your AI until you trust your compiler. This rule will be true for decades into the future, I'm willing to wager...
But, ultimately, software has to follow certain rules, or it just doesn't work. Proper workflows - involving review - are needed. Because security is pretty much over, otherwise.
Folks are finding it easier to make their own software now, too - rather than use others. I predict an end to the app stores - or at least, the primary interface is going to end up being "describe the app you want to use today" instead of picking words from a list ..
Edit: Since I seem to have touched a nerve - I've been working on a project to solve this: https://www.archme.io if you want to know my thoughts on the right abstraction
I have strong disagreement because it sounds like, by analogy or proxy, we have also "solved writing"
Just to make this clear: if you can define a really good PRD and sophisticated technical specs, and a strong set of tests cases to pass, at the right level of architectural granularity, plus adversarial code review processes that triangulate and weed out most mistakes, SOTA agents can write the code autonomously, at or above the quality level of most human coding teams. I call that "solved" but only if you meet those context requirements. Which is still hard, not solved, at that layer.
Solving writing is not a good analogy IMO. Writing is for human consumption, and cannot be wrapped in objective requirements and verification processes. Certain forms of writing perhaps could be (can't think of one at the moment but I don't doubt some exist), and those forms might be good analogies for being "solvable" or "solved."
I'm not typing keys, but I am very much still concerned about the quality and nature of the code. Coding to me is more than pushing keys
> if you can define a really good PRD and sophisticated technical specs
I still believe we cannot waterfall software, the idea seems like taking a step backwards. How often do we learn about an unforeseen complexity only after getting into the implementation?
In my experience with agents, it's better to be iterative and in-the-loop. Start with a decent description, have them research the code/issue, write up an initial plan/design, work iteratively on writing code and updating design doc, review and finalize the code and markdown. Then future agents will have some resources to shortcut understanding the code base.
Saying that LLMs have "reduced the cost of coding" would be boring. And using your analogy, pencils, typewriters and computers have all reduced the cost of writing, but writers are still around.
The problem is, without PR reviews & strict oversight, we're losing knowledge, system design & control of our codebases & products. Which is why, IMO, the coding is solved but the other parts which used to be so tightly coupled to programming are cropping up as their own issues.
You might still need to nudge the LLM in the right direction or stop it from going off weird tangents, but none of that involves touching actual code yourself.
Ai can push a lot of keys very fast, but not always the right ones
if they need to be reminded to follow the coding standards, visible in the very code they are working on, what has been solved?
Opening the IDE and typing program code. With LLMs you don't have to use an IDE, you don't have look at program code, you don't have to care about coding standards. You ask the chatbot to write you a program/feature/fix and chatbot does it.
Is chatting with a chatbot still "coding"?
The part that isn't fully solved is just the software architecture side of things, do you want library A or library B, or write it all from scratch? LLM can do all three, but if you aren't careful it might go down a route that you don't like. But that again can be fixed with chat, "replace A with B", not coding.
At least that's how I experience it. In the before times each non-trivial code change had a real opportunity cost as it would easily consume two days until I could even estimate whether this is worth looking deeper into.
I think "solved coding" is taking it too far, but for many projects, the mechanical aspect of it has been removed or reduced greatly.
LLMs will have a much harder time "solving writing", because they cannot develop their own style and so are severely limited, creatively. This is less important for coding.
They haven't solved coding.
Programming is an art form. And the better you get at it, the better kinds of ideas (abstractions) you can create.
This is something today's AI cannot do.
If everyone were to permanently switch to AI for software development, software innovation would cease.
For example, nobody on our team writes manual code anymore, we have basically set up a harness where an engineer types up the requirements for a change, the system implements it given certain constraints, we have automated unit and integration tests that are ran, and if any errors pop up, they get fed back into the loop until fixed.
But to do that, you need to actually know what you are doing - you have to have good instructions to keep the agents in check and not start making mods outside of their bounds especially when the issue is with a dependant service that is causing errors.
To solve something, there must be a defined problem, what is the problem that was solved. Or perhaps it is just "coding is solved" is the turn of phrase de jour be ause we haven't yet found a more succinct and accurate way to describe the paradigm shift
When it comes to non technical people using Ai to build things on code, the outcomes are on average pretty poor, which i see as evidence that the driver and their expertise behind the Ai matters a lot. A notable example is Terence Tao's conversation with ChatGPT, us math normies could never have done that. The same applies to coding agents ime
Except a little worse, since they were raised alone in a library, act mostly the same, and have harsh limits on personal growth.
My team recently spent two weeks on a wild goose chase trying to figure out why TensorFlow Lite was generating nonsensical OpenCL kernels. Well it turns out that LLVM had a few bugs in the RISC-V assembly for our platform that was leading to silent garbage. It took combing through assembly dumps, hexdumps, a lot of pain staking debugging, and going through the TensorFlow Lite source code to to track this down.
In your opinion, if code is the wrong abstraction to be working at, how do you approach this scenario?
To your point, it's not the wrong abstraction for solving code level bugs. Just like python is not the right abstraction for solving memory corruption or pointer mis-alignments.
Which is really the same problem with coding.
The agentic model of it just taking over and doing everything is poisonous to effective long term team work.
We're well past the point where it's about the quality of the work they produce. It's the way they integrate (or rather, don't) into human practices.
This time I add another definition "when you can own it".
I think this article misses the mark in a few ways, but this is the biggest. I see AI as the ultimate solution for maintenance, in three ways:
* AI does a great job of refactoring and so tech debt becomes shallow. If something is not architected right, it can be fixed much more easily than ever before. * Bug fixing is also a great AI strength. In the future there won't be a backlog of all the bugs that never got fixed. AI can fix them as quickly as they come in. * AI doesn't get bored, doesn't get tired, and doesn't care how crufty the code is. It is happen to maintain any application without judgement.
If you can really get a good set of requirements, go and write all of your test cases out, and then throw it at an AI that will one shot it. Its perfect.
I've tried this out. Even with relatively small applications and with spending hours reviewing the spec documents, there was always something I missed or something that wasn't quite right when seeing it live.
But after 2 months it starts to backfire me. I still know nothing. I have some understanding of the system design and core components but I have zero clue about how certain things are done under the hood. Because AI read code for me and code for me and I take it as my own understanding.
In last week I end up limiting my AI usage and forcing myself (it is really hard) to read and code at least a bit by myself to start having any idea about what is going on here.
Intelligence is the new currency, but it will be short lived
Here's an anecdote about a way to do this wrong.
I have found with AI coding methods that there's a line where it becomes a hail-mary (in the American Football sense).
A hail-mary is when you throw the ball to the end zone and just pray someone catches it. This is almost always at the end of the game.
This moment with AI code is indicative that you can't put together a coherent plan so you just tell the agent to "make it good". It used to be that the results here would suck, but now the agents are really competent, so the results might be good.
But at that moment, that's your cue to back up. Because as soon as you take a solution that's so far detached from your understanding, you're underwater. The hail mary is not part of a larger game plan. It's the last play of the game. There's nothing after.
So as soon as you reach that moment in your coding, you're signaling that you're done understanding not just the code, but even the way it works at a high level. If you're still going to work with this code after, then back up and work with the AI to get more understanding of the problem.
So that's why that Andy Weir's book was named that!
You can still know things and get force multiplication out of LLMs, if you are disciplined and caring enough. In practice, most people won't be. And you can't force other people to be. But you can force yourself to be.
> You can still know things and get force multiplication out of StackOverflow, if you are disciplined and caring enough. In practice, most people won't be. And you can't force other people to be. But you can force yourself to be.
Aside: this reminds me of running a “negative split” in a long distance race, where you aim to run the second half faster than the first. It’s very hard to do this because you have to be willing to let everyone else in your pace group pull waaay ahead, and running above race pace in the beginning feels “free” with all of the adrenaline. But if you do manage to stay disciplined, it’s a fantastic feeling to reach half-way with gas in the tank, and then start to reel in all those runners who sped by in the beginning.
well as you said, the market will decide in the end. You may be even further behind in years 2 and 3. btw, we're already in "year 2 or 3" territory for some, i wonder how those companies are doing vs their competitors who adopted AI full throttle?
Of course things can change drastically in the next years, as they already have in the last few.
Only thing I know is that it is really stressful working on this rat race
And of course, this is very much dependent on the switching costs. Which sucks, because tech companies are great at making this as high as possible; the EU is trying to do something about this with their data portability legislation.
Anecdote: I worked at a company providing infrastructure as service. We lost some bids the first time around, likely for a variety of reasons (price, features, ...). And, some customers came "back" to us after using our competitors' offerings and being burned by their reliability. Yes, switching costs were real, but the pain of losing their customers was even higher.
We have both the tools and the skills.
Later, we will loose the skills because of AI and the depletion of natural resources will lead to the scarcity of the tools.
Dealing with legacy messes I used to get frustrated and bored of making the improvements, now it's easy to clean up code bases and write loads of tests. I asked Astra to come up with a plan for Playwright testing the whole App, I have not built it yet but the flows suggested were fantastic as was the ephemeral database we plan to create for CI.
I built my friends portfolio website almost entirely vibe coded in 4 hours and it looks unbelievable, we added so much slickness (he's a designer) just prompting together. I used a CMS I had never used once before and it was so so easy to do without any of the usual need to read docs about everything.
I've done so much devops now I'm actually fairly confident that me and an AI can do anything you want in terms of deployment/infra and scaling from AWS to Terraform to whatever.
Anyway my main concern about this technology is not that it is crap at coding it's that the improvements in how it codes and thinks are absolutely dramatic which is extremely scary - it has come so far in a year I wonder what the next year will bring.
While these criticisms were technically true, they were stated mostly out of a sense of insecurity from people whose jobs were basically spending years and years just glueing code together and mixing APIs to display some CRUD apps rather than out of a genuine concern for whether LLMs were actually producing poor products.
Nowadays those concrete criticisms don't really work anymore, LLMs are pretty good now and surpass most developers when it comes to writing the majority of shovelware that people have been employed, and so the narrative is changing from concrete criticisms about how LLMs were genuinely not capable of writing software... to these kinds of abstract and philosophical arguments that are really hard to argue against because they make no concrete claims.
If you say an LLM can't implement a feature, we'll we can test that claim concretely and LLMs are getting much better with every new release. If you say its code is slower, buggier, or less maintainable, those claims too can be measured and once again they're getting really good at these. If you say it takes longer to complete a task or requires more human intervention, we can compare it and measure etc...
But now the objection has shifted not to LLMs are incapable, but people are now incapable and LLMs represent a degradation of the "craft". And here there is nothing left to test or falsify. The argument has stopped being about whether the LLMs work, because that's verifiable and they are now at a point where it's hard to argue against their ability to actually produce functioning software, so now the argument is about whether people are morally, culturally, or intellectually permitted to use it.
Whatever the path is, regular engineers (90% of the people around here) will get screwed up one way or another. But hey, playing with LLM agents is cool!
Are you suggesting not using the tools? A kind of technical version of the Amish way of life?
I dunno I think my guidance and testing is very important and I make sure not to ship things with bugs and code that is really awful. I'm still just about necessary for now.
Do you need all these layers of abstraction when the human is no longer looking at the code?
Best analogy is forgetting how to use a slide rule following the advent of calculators. The former was made to make hand-calculation of logarithms easy. The latter does these calculations directly (obviating the need for a slide rule at all).
I think what humans still need to learn are the theory and domain fundamentals for their industry. If that industry is computer science, that means algorithms, calculus, linear algebra, etc. I think the future of CS is then (a) theoretical human-drive design and (b) prompt engineering to implement and verify that design.
It would also be helpful to have domain knowledge outside of CS as having the skills to build something is nearly commoditized (outside of the above fundamentals).
Even hygenic macros still end up being there primarily to increase legibility and ergonomics. Good ones, like core.async, make it easy to understand how threads and the like are glued together, but ultimately it's just syntax-sugar on steroids.
If the goal is not for humans to read the code at all, I'm not entirely sure of the point of syntax macros; the LLM could just generate the expanded code.
The architecture and the intent don't matter to business folks, it never has, and with LLMs it matters less than ever.
Vibe code into production is satisfactory regardless of architectural understanding or intent. If there are issues, just have the LLM spin up some agents to play whackamole until the issues are pushed beyond visibility.
Ultimately, the idea that we can hang onto fleeting engineering disciplines misunderstands where the industry is going, regardless of any assessment of the LLMs capabilities.
Scheduled a quick call to align me on what he expects - normally he wouldn't do that but he has attached a big agenda written by Claude what the presentation could show, and invited two other product colleagues of mine.
I came to a Miro board of the Claude Agenda, put into Miro using the MCP.
Honestly, just tiring. Asked my colleagues if they would just put the Claude agenda into Miro with Claude, what they need me for when an AI could just narrate it.
We all interact with systems through mental models, but if many devs are just prompting claude when something doesn't work, they might read what claude found, but lose out on the exploration, debugging, and work that builds and reinforces the correct mental model and discourages the wrong one. And if devs are missing out on the mental models, will they actually be capable of driving efficient solutions to problems as the mental models get worse.
Does anyone single person understand what’s happening when compiling a large C++ code base? Meaning, can anyone track the basket cast C++ language constructs down to the Clang IR to the optimized machine instructions? From there can any one person follow those machine instructions all the way through to the actual registers etc to actually running the code?
So, if you're pooping out code, and committing it because tests still pass, and that's all you know, you're in for a treat. When an executive wants to know why a b0rked feature lost their department millions of dollars, guess who will have to answer for it, and its not the LLM.
My advice is to find ways to keep on top of how it all works, and if you're the only one who cares, well, then, that makes you even more valuable, not less.
This is also freeing up time to explore things that before you wouldn't have been able to even start. New fields in tech are opening up. It's all about the model, compute, plugins, third parties... and more to come!
And this is not limited to code, but also how the world works. It's all being abstracted into prompts. Funnily enough, I'm also learning that way...just taking less time to get to the point. But this is not the first time we go through this. Eg. Google vs a library. And like anything, if no one knows anything anymore how do we distinguish from one another? There's a level of wanting to understand in order to distinguish ourselves from the rest in the serendipity of everyday life.
Call me naive, but something tells me we're going to start being much more open to just exploring the world with all this time we just bought ourselves thanks to technology. We were always gonna get to this point and there'll undeniably be bumps ahead.
Enough, already.
[1] https://github.com/jackyzha0/quartz [2] https://github.com/sspaeti/second-brain-public
> You may not write the code by hand but you understand it enough to investigate and fix it when it fails. It is how I think we should leverage AI instead of becoming a meat proxy.
[0]: https://raahelbaig.com/entry/responsible-human-in-the-loop/
Not only that, but even if we assume best effort on the engineer, the business pressures don't often allow that. I'm under constant pressure to deliver more, faster with less people. The performance eval ladder at my company was just reworked to double the amount of deliverable features expected per job level per year. Our CEO told us we should be able to deliver what took us the past decade to deliver in a quarter, every quarter going forward.
How could you possibly have a human anywhere near that loop with those demands?
Even with unlimited spend, it seems immensely beneficial to dig into the code base and fix a certain amount of bugs oneself. Oftentimes this is ends up being quicker than having to type out a detailed explanation of an issue in plain language, with the added benefit of maintaining intimate knowledge of the code base.
It absolutely won't be long until product managers are the only humans who actually need to be involves with the software development process.
Tail risks have always existed in software development. The tail risk of a bug introduced by some dev who quit five years ago is similar to the tail risk of a bug introduced by Claude six months ago. Deal with it by building better visibility into how your systems work. Demand that your agents write good documentation to accompany their code-writing.
If you’re doing it right these days, it means you’re thinking of a much bigger picture and containing downside risks as boldly as you’re expanding the frontier of upside opportunities.
As a person who's a solid generalist with over 30 years in various roles, I am completely and utterly shocked at how little foundational knowledge people in "senior" roles possess across a wide variety of technical fields. I'm often treated like some wizard or oracle for knowing things that everybody in the field used to know, I just haven't retired yet.
AI didn't create this phenomenon, it's just the latest (and probably the fastest) iteration of it. I recall another particularly large iteration happened when Windows NT 4.0 Server saw mass adoption. Suddenly, people who were effectively IT technicians were now sysadmins.
Edit: grammar
Unarguably, the world's technical capabilities have increased by this shift (while decreasing the required technical understanding required of the people managing it). Albeit, with some security implications.
The code change itself doesn't specifically matter. But suffice to say, it was about an AI feature in one of our products.
The code was stamped by Claude driven by a prompt. The prompt was for a ticket generated with the Atlassian AI integration. Atlassian had digested docs made with AI. The docs came from strategy memos I'm 90% sure were written entirely by Claude.
The strategy was chosen by management at the urging of exec leadership. The execs now communicate mostly via AI written memos. I do not know how they make decisions, but they reference tech influencers, market conditions, customer expectations.
This gave me pause. Who had actually made the decision then? Arguably there has been several layers of human review, but the actual source of the decision was hard to pin down.
We were not building the feature because we wanted it. We were building it because we thought other people expected it.
Perhaps reflecting on the state of the market, I thought, could indicate who was actually in control.
Where do investor and customer expectations come from in 2026? It is very murky, at least in tech. There appears to be hype. Some hype comes from true believers, some comes from cynics. But both respond to market incentives that reward bigger and bigger claims.
Where does the market's "action" come from? What is the driver?
Investors do not really seem to understand what the tech is or its limitations. Some are passive operators. Others are just responding to the overall froth and speculation in the market - which becomes a runaway feedback cycle.
This left me lost.
Nobody in this ecosystem, I thought, is actually in control here.
Nobody is actually orienting work and action to real, concrete goals. It's all based on speculation and anxiety about the future.
So it is not only that nobody understands what the code does. It is that we cannot, or at least I cannot, explain the motivation. There doesn't seem to "be" any form of "intention" in this environment.
It has all been hollowed out, replaced either be inscrutable machines, or inscrutable incentives.
Ironically it rather resembles the kind of "misaligned" superintelligence we are supposed to be avoiding.
Data can describe to you what exists. But it can’t tell you what you value.
What you describe is people who can’t tell the difference, and who let the machine (data) make the value judgments.
This is a long winded way of saying "people made up clever/sensible sounding stuff". Now it's easier to do it with AI so the problem is worse. However, I'm not sure what you were looking for was ever really there - the "inscrutable machines" and "inscrutable incentives" were always quite inscrutable.
The hyperbole feels like it's been contrived to fit into the classic AI counterpoint: But humans do this too!
It bothers me to no end when I get an AI written response, especially from the executive team or any one of my co-workers
Was there never a developer in the loop? I hope my org won't give up control of their business to an AI like this soon, sounds like a nightmare to figure out what's going on.
Who's in control? Everyone is, to some extent. And no one is: when you're hungry for food, are "you" in control of that? You can consciously repress your impulses to go eat something, but your mind didn't create those impulses.
Human societies develop impulses and minds of their own, emerging (weakly) from the impulses and minds that comprise them, and they make decisions in mysterious ways.
Of course, it sure is nice when we can come up with a compelling story for the motivations behind something. Easier said than done…