Posted by drogus 7 hours ago
You can do this in a professional way but still get the point across that they're stealing productivity from others that have to pull up their slack.
If the commits are too big or too dense. Make them break it up, clean up the comments etc. Otherwise they are simply doing a poor job. Not Claude, them.
Can’t have the cake and eat it too.
If they “solve” navier stokes and mathematics, they sure as should be able figure such menial tasks out on their own, shouldn’t they?
But it’s the fat part of the bell curve. There’s loads of even worse crap outside corporate code…and at the right end of the bell curve is deeply thought out and the result of long-term maintenance that tends towards as bulletproof as can be possible in this universe, parts of OSes, networking stacks, clocks, and the like.
It’s only because that stuff is so robust and sufficiently general that the rest of the world’s steaming pile of code works at all.
One of the worst things you can do is pull punches when you get "contributions" that are net-negative because they waste everyone's time.
It takes effort to debunk because you have to holistically consider what the problem was and actually find the better solution to prove why it’s lacking.
In other words, to debunk it, you have to do the actual work that wasn’t done the first time.
Some people may argue “so what?”
And to that I would just respond that the fast solution implies nobody probably thought about it, which is always risky and generally leads to very bad outcomes. If for no other reason than there is at least no consensus, which for long-term software evolution is deeply problematic.
This has happened many times since this trend has started, and forcing everyone to step back after a year objectively reveals the murder that has been committed that, believe it or not, is not trivially unwound.
In the wrong hands it’s a debt machine that is already killing companies from the inside.
Then respond with that addressed, with a firm but polite explanation that the document has some problems that need to be addressed carefully before we invest time on it.
None of this works if it’s coming from your boss, but it’s very effective at making people think twice before sending you slop.
People only do the workslop thing when they believe the benefits of showing the work outweigh the reputational risks. If they get caught every time they do it and exposed on an email chain, they start doing less.
BUT because it is AI it's not that section that's addressed, it's about 90% of the document that now has changed, and I need to spend another gargantuan effort to review it. It's like by trying to be helpful I actually lose more time.
The AI has made it so that the thinking before writing is mostly gone, and shifted that step to the reviewers.
Now, no matter the amount of preparation work, those last rounds can completely rewrite something with no hope of review. No iterative improvement. No ratcheting towards a known quality. Just a bunch of cargo cult review followed by YOLO-style, vibe-everything absurdity.
The ratio of verification capacity to generation capacity, V/G, has broken with LLMs. It’s not simply an issue of more generation or less review.
The impression seems to be that individuals are more productive, but that productivity is someone else’s review burden. So the team/firm as a whole is not better off.
The cheap generation of content does mean that reviewer capacity is now a limited resource.
Unless your firm is aware and is measuring time spent on reviewing slop, there is no incentive or structure to ensure that time is respected and valued.
This is a management and awareness problem since the typical response is “use a bot to review it.”
Not to say these proofs are slop, even if the LLM-generated work is of good quality, what happens when no one knows how it works anymore ? Even if you ask the LLM to explain, which it does quite badly, the time to understand the explanation is incompressible.
So in the end, productivity will probably reach a ceiling that we can estimate as the product of humans, their cognitive capacity and their time. And that ceiling may be lower than what AI companies valuation expect, regardless of the compute and they can pump out and the RSI level they can reach.
Presentation layer code doesn't control how the machine and kernel prioritize anything; so "proof" Ruby code is doing the right thing is proving the machine does the right thing from the factory.
As for abstract theory and math, well shit since any English and any math are...mathematically possible...well shit I guess we gonna have to live in the real world and not inside a rhetorical bubble; religious or atheist philosophy... cause they are not evenly distributed frameworks as religion clearly shows; so why live by the syntax and semantics of some mathematical rando who taught a stats class years ago?
Same shit as living by religious allegory
Goodhart's Law has come for 1900s means of scientific inquiry; every technology follows an S-curve and the same for every social society. In the US we aren't all defaulting to calling ourselves British or speaking Latin.
Physics will always be there. The stupid glyphs and bird song we came up with to communicate about it isn't physics. It's just a human language system.
if your prompt is not specific enough, many decisions are a coin flip.
If you are using AI, you still have to know what you want it to do. If you are a project manager and give your team incomplete requirements, the result may not be unlike this.
In some cases you can get away with not reading the code. And maybe in the future that will be more common. But for now, I personally prefer to read (or skim) the code, and not abdicate to AI.
If I'm tired, yep I'm going to write more sloppy code, I'll put less thought into it and be more tempted to tell the AI 'just do it'. That hasn't changed. I'm not sure where this new breed of "don't read the code" is coming from when it's actual software engineers making the statement; maybe it's just over-excitement and it'll die down (hopefully).
Or it's just over-simplification -- I don't read ALL the code either anymore, but at the same time I'm reading SOME of the code. Tell your AI agent you're doing a code review workshop, ask it to select ten files at random and then go through them with you one at a time. Ask it to review common anti-patterns and put them in your coding standard. Then you can start asking it to review and fix code against your standard as a first pass. You'll still need to manually review the code to keep it in check, but at least you can make it a bit less daunting and a little more fun by actively collaborating on the AI with the review.
I'm learning it's just a 'different kind of tired'. AI has made some stuff much more efficient, including the choice between efficiency and carefulness. You can make that choice at a micro level and very quickly now if you want. There's still fatigue though, it's just a different kind of fatigue. You still get tired from all the thinking. You can still choose to put the work in or not and it still makes a difference. It's just a different kind of difficult. Doesn't mean that the thing to do is to throw your hands in the air and declare that your job shouldn't be difficult or include hard work anymore.
If we are going to the moon, then by all means I think we should probably scrutinize every single line of code, but most business aren't going to the moon.
He favorite axe to grind was how my generation put too much trust into the compiler. "Just because it compiles doesn't mean you're done. You need to look at the actual assembly. Compilers can be VERY inefficient."
This was in the 90s. There was some merit to his claim.
But today? How many people who write Javascript look at the actual assembly code that is run? Not many.
I can't help but wonder if this current "you need to understand the code" is the same thing all over again. And if in a few years, almost nobody will look at the generated Python, Ruby, etc.
They don't look but they should.
I think the main problem is that most people pushing for heavy AI delegation either
1. Don’t really care about system reliability or 2. Don’t understand the difference between code and system, and assume that a system must be reliable if some automated checks pass
remember this timeless Agent Smith quote
With UBI it won't be expensive anymore. So much of the economy is simply extractive, most people's jobs are really already highly algorithmic and involve much less "thinking" than they think. Being able to sit around and think has often been a historical luxury - of the elite or those lucky enough to be subsidized by them. I suppose the common (as they all were) man of prehistory also had more time to think, which was doubtless the germ of humanity's religious and speculative impulses. But they also had to contend with a brutal world where death was around every corner.
The economy doesn't want you to think. Knowledge workers really have far too high an opinion of themselves in this regard. Your "thinking" was merely more instrumentally useful than the alternatives. Most of us have done very well while inventing nothing. But an even better day dawns.
When humans don't have to rely on their own labor to survive, more humans will think. The opportunity to afford to indulge your curiosity as if you were among the wealthy and privileged. A society that can afford the Enlightenment and its myriad avenues at scale. What else will there be to do?
If you are not using AI agents for everything, you're doing something wrong, and wasting company time. It has been emphasized that no one should be writing code by hand and if you think an AI has implemented something wrong you need solid reasoning to show why or else people just label you as some anti-AI troublemaker, who just becomes an obstacle standing in the way of things getting done.
If your human generated opinion is in any way wrong or simply not a true homerun then you get snubbed in future reviews, people stop listening to you.
I'm going to read you charitably, and assume that what you really meant is "the burden of proof is entirely on you, and is set unreasonably high." Because of course, if you do think someone has done something wrong and want to say it publicly, you do need a solid reason.
However, it doesn't take away the concern of architecture, design, etc. and questioning if the current solution is well designed or not.
In other words, if you can't question the AI and refute it in a topic, you aren't expert enough to use AI in that area in a engineering manner.
Granted, not everything needs an engineer behind it. A shack to store some tools will survive long enough probably
They argued this still allowed them to fully understand what has being implemented and how it fitted together.
I’ve been doing this since I read that and it also allows you to catch stupid stuff while your typing, you can reason about what each little change does and why it’s needed.
This also lets the LLM change the future of the plan if you fine something.
This turns out to save a whole bunch of time later because you already know how it works.
It’s not nearly as fast as just letting the agent do everything.
I've also been thinking about the partner programming craze phase our industry went through. I write the test, you write the code; well I write the test, LLM satisfies with code. (Then we have less of these weird LLM generated test cases that test _nothing_. We can also use the LLM to suggest tests to complete coverage.)
These would be slower to work this way, but the end result is:
1. humans still learn 2. you have a good grounding in how everything has been written
That is if I ignore the slop MRs with 60+ files changed that I have to painstakingly go through and then politely tell the author to fix the crater sized holes in it, while a vein in my forehead almost explodes. And some part of my sanity is lost forever.
If it were me, I would start by getting a summary of what problems the existing solution solves. I would also make sure that summary had the non-functional requirement. Maybe write a test suite that opens the page and does all the stuff.
Then ask the LLM to implement in another language, with those solutions in mind, but first coming up with reasons why the target language might not be as good, and where it might have useful features that the original language didn't. Use the test suite to check if things are working.
I honestly don't think you would land far away. You would have to answer a few questions along the way, but it would mostly be plain sailing.
I would expect to be able to do this with very little human attention, whereas a language rewrite two years ago might take a whole month, not including learning the new language.