Top
Best
New

Posted by dimonomid 13 hours ago

AI Has No Wisdom and Neither Will You(alexn.org)
366 points | 510 commentspage 2
davidee 13 hours ago|
There are already consequences.

We (collectively) were unprepared for a machine that presents itself in human forms. We were the frogs that boiled ourselves. We built a world of images and words on a screen. And then we built a machine that can (increasingly) mirror that world; it does so in a way which most of us are incapable of disambiguating.

It feels like there is indeed a ghost in the machine.

And there is, but that ghost is us. And that ghost is fading surprisingly quickly.

addag 13 hours ago||
As much as I'd love to believe it, it is now a conservative take. Sure, having a solid architecture in mind still matters right now, but manually writing code is completely unnecessary and, before long, even designing the architecture is going to be completely automated.
tripledry 12 hours ago||
Here I am starting to think a central part of knowledge work is the knowledge gained.

Maybe you are right, maybe not, let's see. Unfortunately I kinda agree, because humans are really good at being lazy and going in the path of least resistance (including me). It's genuinely difficult to not use AI even if it makes my work worse, as long as it's easier and faster.

datsci_est_2015 13 hours ago|||
Does writing table schemas count as writing code? Is that architecting? Both?

I personally don’t trust coding agents to have enough context to write domain-specific table schemas, and I don’t have the patience to transcribe all of the context into a natural language prompt. If I ask it to, it’ll write something for sure, and maybe that can be a jumping point for me, but at some point I have to physically write what the columns will be.

altcognito 12 hours ago|||
You should definitely try it if you haven't already. Most of the models are deeply familiar with domain specific areas in ways that most of us aren't. In fact, I would say the problem is generally the opposite. If model size and effort is is high, it will over architect what an app needs. I find myself reigning in a giant email notification system with subscriptions and "channels" and other such nonsense when sometimes you just want a simple one off.
datsci_est_2015 12 hours ago||
Yes, the LLM knows the definitions of the words in my domain, that’s clear and obvious.

No, the LLM doesn’t know our product strategy and why certain things matter and certain things don’t. That’s very much what I get paid to do. There’s not even agreement within our team about what path we should take through the domain-product space, there’s no chance in hell that the LLM will choose a profitable random walk through that domain-product space.

If it were able to do that, then AGI would have already been achieved and we’re only compute power away from OpenAI or Anthropic making the marginal utility of any piece of code $0.

altcognito 11 hours ago||
Yeah, it's not just over-engineering, it is just that generic systems that look like everything else usually don't solve a business problem, but I think I would be less dismissive of its world model in general.

LLMs are not a "random walk", they take in information and they explore the space according to the way they've been instructed.

addag 12 hours ago|||
I repeat, I hope you are right as I enjoy coding a lot. I just wish that people who keep not using LLMs too much won't have troubles because of that.
ttul 12 hours ago|||
I’m taking this view now as well. If you’re reading code, you’re probably doing it wrong. You should absolutely be setting criteria that can be objectively measured and rejecting code that doesn’t meet those criteria or perform as specified. We are all senior software engineering managers now, with a fleet of cheap and ambitious young engineers doing all the authoring.

But reading code? What does that accomplish, other than to slow your dev process down enormously? Serious question.

NichoPaolucci 7 hours ago|||
Man, I think that code is still the artifact that we produce as developers. Code is the truth. I don't find it difficult or super time consuming to just... read the code, either. I've highlighted quite a few issues with LLM/Agent output from just glancing at the code.

I'll let you know how it goes... My new VP of engineering is a 'no looking at code' type of guy and is ripping 10K LOC PRs / Docs / plans against our 25 year old codebase and I would not say that they're 'good' PRs.

Maybe I'm completely wrong, but I think reading the code is more valuable than ever when working in a full-stack / small company role. I can tell you exactly what the business logic or functionality is for a certain piece of our system, in truth, without having to step through and make sense of ambiguous docs (that were also AI generated).

(I have a sneaking suspicion that in two years or less, my small team is going to significantly compromise the integrity of this codebase. Maybe by then we can refactor with GPT 12.)

kkapelon 12 hours ago||||
> You should absolutely be setting criteria that can be objectively measured and rejecting code that doesn’t meet those criteria or perform as specified

This is the classic "make no mistakes".

On a serious note, I might set as criteria "avoid code duplication". Does that mean that the model/agent will actually follow it?

> What does that accomplish, other than to slow your dev process down enormously?

I am an OSS developer and I often see PRs (i.e. from the general public) that look correct, pass all CI checks, are heavily documented and they are still wrong.

Most of the times either they duplicate code that already exists somewhere else, or they implement a "feature" by opening a can of worms for subsequent "features" in the same area.

yurish 8 hours ago|||
How do you measure absence of concurrency bugs for instance?
y-curious 13 hours ago||
But then what will I do for work? :(
addag 12 hours ago||
Our only hope is that work will cease to exist.
pixl97 12 hours ago|||
[Looks at people running AI companies]

Hmm, I have deep concerns about what and who is going to cease to exist.

addag 11 hours ago||
Well technically my take is still valid
elric 11 hours ago|||
And we will all live happily ever after? What a load of naive nonsense.
hypfer 13 hours ago||
Tbf, at this point, this has been said ad-nauseam.

At least I did not find a new thought in that (granted, relatable) rant.

"This is bad and you are bad" requires people to not defend their reality through rationalization, but the point we're at with AI right now is driven by exactly that. So this is at best highly ineffective at reaching the people it claims to want to reach.

That said, the underlying emotion of "you all suck and I hope you lose your jobs you frauds" is relatable and worth screaming from the rooftops of Linkedin dot com for the catharsis alone.

itomato 13 hours ago||
If you have ever looked at a kite and said to yourself, "that is as good as a hang-glider. maybe better.", you might be subject to the perils of one-shot AISDLC
grim_io 11 hours ago||
Every large code base goes to shit unless you have a very strict bdfl at the top.

AI is nothing special or new in this regard. It just gives a single guy the velocity to ruin a codebase at the rate of a full enterprise team at double the speed.

A shitty code base still makes money, and that's all that will ever count for the majority of the employed developers, those who don't blog.

FinnLobsien 13 hours ago||
To me, the problem is not about what AI is good or bad at, but about institutional knowledge. If AI is doing the coding, writing, designing, or whatever else, you slowly lose the ability to a) learn from others in the org because nobody knows what you need to learn anymore and b) actually improve stuff because there's no more "what good looks like"
weego 12 hours ago||
My companies current national marketing campaign strategy, structure and creative work has been created by an MBA from outside our industry who did a "deep dive" using AI. The project manager they brought in to run the tickets is also adding AI made decisions on brand and ad copy.

The short version of the result is that we're completely copying the leaders in our market that have 100x our marketing budget because they're the only ones with documented strategies for AI to cheatsheet from, with no sense or irony or concern that perhaps their scale is a core part of their strategy.

It's not just institutional knowledge. We're suffering the death of the specialist. Now every generalist is pulling triggers with no sense of limitations.

matthewmccc 12 hours ago||
which begs the question - will institutional knowledge not matter before long?
mschuster91 12 hours ago||
it will not. Dead Internet Theory - eventually everything will be slop. And even if you care to avoid the slop, there is no place left for anyone to differentiate from the market slop offering because the 10-20% that actually care about not being fed slop are not enough to sustain a market.
pixl97 11 hours ago||
Really it seems more like late stage capitalism and things that were already occurring before gen AI in all markets, not just internet based things.
alexghr 10 hours ago||
I think (hope?) our industry will reach some equilibrium between fully agentic development and reviewing/hand tuning the code.

Personally for my job I struggle to justify writing code by hand. Agents are much better at typing the code, reading the code, finding patterns and discovering bugs. They don't have my knowledge and my background so the current models still miss things or overcomplicate the code. I don't know if that's going to be true of the next generation of models (or the one after that).

With current models, I cannot fully give in to the vibes. I _have_ to look at code (maybe not 100% of it but at least a majority of the lines). I _have_ to grill the agent to make sure it writes the best code that I deem possible for the problem at hand. This isn't the fastest way to do agentic development, but AI agents have already made us significantly faster than we were in the pre-AI era. I don't want to squeeze every ounce of 'development speed' if that comes at a cost to code quality, maintainability or debuggability.

For my side projects, things are different: either I only care about the outcome (and I give in to the 'vibes') or I care about the journey just as much as the destination in which case the agents are extremely advanced search engines, code reviewers and mentors but they don't get to write any of the code.

Yeah, I guess I personally fall on the whole spectrum: for side projects I'm at the extremes but for my professional work I fall somewhere in the middle.

sobiolite 13 hours ago||
> Fact is, vibe-coded projects devolve over time into an unmaintainable mess.

We've had strong coding agents for less than a year. Anyone making such a definitive statement about how vibe-coded projects progress over time is basing it on guesswork, not evidence.

abirch 12 hours ago||
In my experience, the most difficult part of a software project is getting the specifications correct. To abuse Harold Abelson famously quote: "Programs must be written for people to read, and only incidentally for machines to execute."

"A Project must have proper tests and specs, and only incidentally for a working program that executes"

gedy 12 hours ago|||
Somewhat, but I've for sure seen new projects become grindy messes in a few weeks.
mohamedkoubaa 13 hours ago||
Or, evidence from a sample of poorly designed projects.
AnimalMuppet 11 hours ago||
So there's evidence that it can be done badly enough that the problems show up in less than a year. That doesn't prove that more evidence will show up after a couple years, or five, but it does make one suspect that it will.
mohamedkoubaa 4 hours ago||
Stewardship requires a sustained commitment to quality and in some cases people just care less over time
tylerjharden 12 hours ago||
As far as maintainability is concerned, have we all forgotten the 12 Factor App in the age of AI?

https://12factor.net/ https://en.wikipedia.org/wiki/Twelve-Factor_App_methodology

Kevin Hoffman expanded on that with the 15-factor app: https://developer.ibm.com/articles/15-factor-applications/

Mind you I may be dating myself as I was first introduced to this paradigm in 2015 working as a Java SpringBoot engineer on an enterprise project that I then migrated (57 microservices) all to Scala, after onboarding two weeks to Scala fresh from no prior Java experience.

I feel like there is so much "wisdom" encoded in books and writings from some of the most prolific engineers and architects over the last several decades.

Look at Matt Pocock's skills with simple primitives like grilling the human, researching through wayfinder maps (a Godsend to my workflow prior to Cursor Projects and orchestrator patterns), and having a solid Domain Driven Design through defining a shared glossary and breaking up work around proper seams.

kkapelon 12 hours ago|
12 factor is good but it is a very low bar.

It is perfectly possible to vibe-code a badly designed app that still passes those 12, 15 or whatever points you define.

tylerjharden 11 hours ago||
Could you share some higher bars that may make up a better "rubric" for this issue? I am actively trying to do so. Shy of just condensing core Manning publications that cover domains of interest, I am struggling to find a good bar to have my clankers validate against outside of minimizing cyclomatic complexity.
kkapelon 10 hours ago||
Cyclomatic (and cognitive complexity) are a good start.

Some other ideas

1) Enforce architecture decisions (see archunit). But somebody needs to write them down first.

2) Check that tests actually break if the code that accompanies them is removed (several LLMs/agents today create tests that don't actually test the code they "guard against)

3) Automated performance testing. An LLM/agent might create a change that is "correct" but increases latency for 3x (best case) and 20x (worst case)

The hardest part that I see no solution for today is to understand when a change breaks backwards compatibility. LLMs/agents are trigger-happy and will happily refactor/remove stuff without any care about who is using that.I don't have a proposal for that, but the problem is there and is not covered by 12-factor config.

tylerjharden 9 hours ago||
Excellent set of criteria, thank you. Something like am archlint, or performance lint, and sanity check on tests "Actually testing a seam or function" of the actual code base sound like good research avenues.
jmartrican 11 hours ago|
Solid points. I will add a few points.

1) A lot of the time i spent deciding on interfaces (methods, classes, etc.) for humans. E.g. should this be two methods or one, should this method be in this class or moved to utility. Those problems went away. 2) What about performant code? This can be prompted away and when the measurements in your performance tests do not go down, then you can step in. 3) Sad to say but the AI has always been better than me at code-reviews. Maybe this is just me and if so I own that, but to the articles point, it might be harder to fix now. 4) "vibe-coded projects devolve over time into an unmaintainable mess". Preventing and managing this mess is the new skill sets we need to develop as software engineers. 5) Another skill-set we will need to master is how to maintain and grow our coding skills. Some ideas are: a) every once in a while implement a feature yourself. b) no AI Tuesdays! c) Have the AI quiz you on the code base. d) Have the AI develop HTML docs about how the code works.

__MatrixMan__ 11 hours ago||
I've found it helpful to make an explicit delineation in the codebase for which interfaces I'm going to agonize over and refine and which I'm going to ignore.

One set goes into haxe files and I use reflaxe macros to write tiny compilers that generate docs, clients, servers, cli's, test cases, serializers and deserializers, etc in whatever language is appropriate. That leaves gaps, which the AI can then fill in.

So I iterate on the haxe stuff. If the AI is struggling to "draw the rest of the owl," I change the source of truth until it has enough guidance re: type related errors, failing tests, and documentation. As requirements change, these things change.

As for the rest of the owl, it's disposable. Every few months I'll delete it and have an AI rewrite it from scratch (now with a smarter model and better docs and API specs and tests to guide it). This keeps the cruft from accumulating while preserving human contact with the code.

NichoPaolucci 7 hours ago|||
> Preventing and managing this mess is the new skill sets we need to develop as software engineers.

How can we prevent and manage the mess if we're not actively working in the codebase though? (Sometimes, I'll get a 'feel' for when something needs changed or will become unmaintainable, but that required consistently interacting with the thing)

Currently, we look at code once during the PR and then we never touch it again until it comes up in the PR.

I'm still handling a few tickets a week without AI, because sometimes it's faster to make a 2 line fix than to write + review a prompt, but increasingly it's just to make sure I still 'got it'.

daveguy 11 hours ago||
1) Those problems did not go away. They only "went away" if you assume no one will ever need to read code again. And if you believe that, you clearly aren't reading the code already. Which I guess is why you think naming and architecture problems went away. AI is the absolute worst at architecture decisions.

2) not even sure what you mean by this

3) probably shouldn't admit that. It implies the reviewer has a less than average understanding of the code they're reviewing.

4) preventing and managing vibe code devolving into a pile of slop requires programmers not use AI. 80% accuracy repeated in more and more layers === more and more failures. In other words, the skill required is exactly the skill of being a good programmer without AI.

5) e) no AI all the time or only use AI as search. You're almost there with a and b. With c, it just doesn't understand well enough to "quiz you". With d, how are you going to know if the docs are correct if you aren't reading the slop?

More comments...