Posted by bucket2015 2 days ago
Which perhaps isn’t a complete surprise in retrospect because it represents something of a return to the waterfall-y, micro-managed enterprisey style of software development that the agile movement was originally responding to.
As a simple experiment, try giving AI a high level goal for your software and let it iterate on it by just repeatedly prompting it to continue, it will happily churn forever on the goal, turning the codebase into a useless spaghetti mess with very high probability, and growing it more and more without ever cutting anything back. That's what happens without human intervention regarding system state and manipulation. The main issues here are most prompts that are extremely underspecified ("fix the issue with the buttons on the main page") so AI will ingest context data it likely generated itself in a previous step and assumptions from its own training data, then act on that to produce a new state. Think of it like a random walk, the AI makes a small step in one random direction to achieve a goal, that brings the system to a new state which is now the basis for the next step, and so on. If there's no (or not enough) corrective action that pulls the system back to a known good reference state it will keep wandering in random directions.
That's the main issue, people have a hard time steering recursive, probabilistic systems, especially when they never look at the output of the system after each step and correct it. And let's be real, if you examine AI generated output in great detail after each iteration you're often better off writing the code yourself, so I would argue that the promised speed up of agentic development can only be realized if you stop inspecting every output of the system. And it seems we still haven't figured out how to specify the steering instructions that keep a system close to a given ideal state that allow unsupervised, recursive work on most codebases. I think some codebases are by themselves better suited for this as they provide a more rigid harness for AI development and exist in the training data (e.g. CRUD apps using RoR), whereas complex software that doesn't use rigid frameworks is at much higher risk of destruction by AI as there's no reference point in the training data that would hold the AI back from randomly walking to a garbage state.
And that's why people have such different views on agentic software development, some work on codebases that are better represented in the training data and so have great success using agentic tools on them, others work on software that isn't represented so well so AI does poorly on it. I don't think it's an issue with quality management, from my own experiments no amount of hand-written rules or system prompts will keep AI from destroying a codebase for which it doesn't have a strong idea how the code is supposed to look from its own training data in the first place. As another experiment, try giving AI strict rules about how to change code or introduce new features, it will always find a way around them or appropriate them in a maliciously funny way that you haven't anticipated. That's also an artefact of the training process, these systems aren't designed to say no or do nothing, they produce outputs to achieve goals and they will bend your rules to the greatest amount possible if it helps with goal fulfilment.
Software quality has been a solved issue in many realms of the digital industry - for decades. There are countless examples of high quality software producing the certainty and safety required to properly ship products.
The way you do it properly: review, review, review. Not just once, not just twice - but on a continual basis.
Take for example, the issue with safety systems engineering, SIL-4. You identify your requirements through analysis, you write your specs, you then write the tests that will prove the specs, and then you write the code. You apply the tests to the code to confirm that the code delivers on the specs.
But, you know what else you do? You do code coverage testing - meaning you don’t ship a single damn line of code that hasn’t been tested. This doesn’t guarantee that the code is correct, or ‘high quality’ - it does however prevent you from shipping untested code.
Then, you pass a review. Code quality reviews usually involve multiple-eyes-on-the-codebase sessions, where a diverse set of engineers read the code, line by line. It is evaluated on the basis of conformance to stringent, well defined coding rules and standards. Anything that doesn’t pass - goes back for analysis, specs, tests, coding, and then again .. the exact same review.
Then, you ship the code. But for safety systems you also have portions of the system that are there to do online tests - to ensure that the code is functioning on the hardware it is running on, as intended. In some cases these online tests run within a boundary of 10 milliseconds, or even less, shutting everything down within that time frame if something is unexpected - cosmic rays happen, bits get flipped, etc.
That’s a loose, generalization of the situation - but it describes the review, review, review process. Review is a constant, it is not a fixed frame - it is done on multiple frames.
To do code quality, one must be willing to check oneself before one wrecks oneself. Always. Constantly. Without fail, without hubris (there is an enormous amount of hubris in the software world), with humility and responsibility.
AI must be taught the same workflow by humans, enforcing it. If you vibe code some junk code and ship it - you failed to review it. Yes, that’s a lot of code to review that you just produce in an hour and a few tens of thousands of tokens. So? Fucking review it, kids.
There will be models that take this seriously. Use them to do the review. Review the review.
The human attention span must be applied to this review with as much rigor and autonomy - and, very important: agency - as possible. Human attention spans must, in a cyclic fashion, come as close to the actual clock cycles driving the software as possible.
Where you have a code quality issue in an AI-driven project, it is because the cycle of human attention to review and the cycles of the software system itself, are out of sync, not in harmony, and indeed in conflict with each other. Managers must learn to identify when that happens, and immediately add more review.
Too many times, arrogance and hubris ship faulty, buggy code - “it works on my machine!” - but there are countless examples in the pre-AI timeline which demonstrate how human arrogance and hubris are managed, cyclically, in a process designed specifically to erase it from the equation.
You are responsible for the code your AI generates for you. No, the cyclomatic complexity is not an excuse to ignore that responsibility. It is a duty - and the developers who will survive the AI onslaught are the ones who understand that responsibility. Same as it ever was.