Posted by jacquesm 10 hours ago
[...]
> I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive. The thing that will work is actually curing cancer. I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world. That is totally on us, and I think it’s the criticism you should be making, instead of all this stuff about messaging and marketing.
> We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months. When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that.
Turns out "winning back trust" doesn't have anything to do with any actual concerns people may have re: employment, electricity prices, stock market bubble, intellectual property, scams, cybersecurity, environmental issues etc. Rather we'll just do all of that even harder and the miracles ("curing cancer", lol) we've so far failed to deliver are bound to arrive in short order!
This is honestly hilarious. Dario must think people are stupid.
First, cancer research charities are some of the most well funded on this planet. They all obtain donations on the basis of "together, we will cure cancer" messaging. They all fund the best science they can.
Second, the human body is a complicated thing. You can feed your fancy LLM as many textbooks and academic papers as you like, but the reality on the hospital ward will always be different. Why do you think student doctors have to spend so many years "doing the rounds" Dario ? They are all academically smart, they are all capable of memorizing text books ... but there is no substitute for seeing and doing the reality.
I barely trust Claude to write code, let alone find a cure to cancer.
It's like saying "we'll cure infectious diseases"
Don't worry, I'm sure that's his next blog post. Right after Claude has finished solving famine and poverty. ;)
A more sophisticated view is that cancer (one idea) is an problem of rates (like most things in biology) and there is a tumor burden per unit time from whatever causes vs. the immune system's capacity to detect and destroy them per unit time. As humans age, the balance tips in one direction. If the balance is too far off for too long, tumors accumulate, doctors take notice, and assign whatever labels to the condition.
Looking into particular causes of tumors (and calling each cause its own cancer) can create a vast research program that employs a lot of people, since there are so many mutations that lead to tumors. But interrupting the development of a particular mutation is not going to have the same society-changing effect that solving the fundamental issue of rates would.
Literally first phrase on Cancer Wikipedia entry.... [1]
"Cancer is a group of diseases involving uncontrolled cell growth typically resulting in tumors with the potential to invade or spread to other parts of the body...."
Maybe I'm overly cynical, but I think charities exist to exist. They don't have a strong incentive to actually deliver on their mission. Everyone at the org may be fully bought it and obviously want to cure cancer, but as an organization it would fail to deliver on the promise. Like most organisms, their actual purpose is to continue to be influential, grow and continue to exist.
I think cancer will be solved by some group that could make a lot of money solving cancer. The probability of any one path working is very low so the payout would have to be very large for anyone to be willing to pursue.
It will become victim of the conflict between its core search for truth, the current administration mandate to spy on its own prompts for approval, and secret orders to hide the mission true purpose, like that other computer we have heard about.
All of those guys think people are stupid. That theme has been repeating over and over again.
I think typical bubble behavior the leaders have set up the whole promise to fail. Everyone is expecting some faux super intelligence to come and find a cancer solution everyone else missed. However, it is just as likely that vanilla current LLM's will create enough of a productivity boost for back office automations in research heavy hospitals to create the space for regular humans to create these breakthroughs, but LLM's won't be able to claim that for themselves and inevitably “fail”.
P.S: I am not saying applications will go away but LLM's are clearly massively effective here at the rote parts of it all.
It's hard to keep track of the frontier on bio ML, but it seems that we're going slower than what Demis Hassabis said in 2024 with 5 years to full cell molecular simulation. We still aren't able to reliably model a tiny surface of the cell membrane.
And of course there's Derek Lowe's takes on the drug discovery pipeline waiting for the proof in the pudding.
To me the only reasonable bullish position is that there is a very non-linear AGI threshold for accelerating progress that we haven't hit yet.
For me personally, I'm looking at other more tractable fields as a proxy to measure this kind of progress. The best modest evidence is from the agentic coding area, (modest because these kinds of gains may not translate to bio progress). Other soft-ish fields to like legal/law/tax are also interesting to watch, as a small amount of people are now trusting AI for these areas that were considered totally unusable a year ago. Another proxy is being able to generate generally entertaining media.
Turns out that “give a reasonable probability of being close enough such that you can bootstrap a solution out of experimental data” gives a very high utility and effectively obsoleted several experimental techniques overnight; pretty much “Molecular replacement” is about the only technique for phasing resolution anyone bothers with any more.
But again, “Bio” is an _extremely_ broad term; for every part of the field Alphafold had a big effect on there are a thousand different parts of the field that it did nothing for.
They made huge progress, but I would say that the vast majority of work on this problem was designing the harness for the model. That's a lot of work for each and every domain.
Isn't that so far only static folding?
[0] https://www.science.org/content/blog-post/so-how-ai-drug-dis...
Capitalism has always been a tension between capitalists, who want maximum return on capital and workers, who want pesky things like a living wage, sick leave or safe work conditions.
For the longest time, the only power labor has had to get those things was the power of collective bargaining. No agreement with your workers meant no production happened.
Now, there's finally a chance at salvation. The capitalist Messiah is AGI and it will finally deliver them from those annoying laborers.
This ideological bent is why they're putting everything they have into AI. It's why VCs and their fellow capitalists are going so crazy.
They see it as a way to finally solve the contradictions of their ideology, but in reality it'd only create a new one: If everyone's out of a job, who will buy their products?
It doesn't matter that I can't do it. They'll pay me lots of money if they think I can.
The underlying purpose of AI is to allow wealth to access skill while removing from the skilled the ability to access wealth. (read on HN, not my words)
Curing cancer sounds insane, but it's also a research problem, not a societal level coordination problem. And one AI has already proved to help with breakthroughs (alphafold). IMO it makes sense for them to shoot for something like that as proof of AI's beneficial sides.
This is not true. There are two companies at the center of AI direction: Anthropic and OpenAI. If there were anyone on the plant who has the ability to influence our direction then it would be Dario Amodei.
Acknowledging the grievances is a good step, but it's not enough. There needs to be a clear explanation of actions to address them, a plan to enact those actions, and commitments with consequences in failure of those actions. Tell people how you're going to make them more employable and effective and needed. Tell people how your datacenters will be carbon neutral. Tell people how financial actions resulting in a frothy market will be coming to an end. He and Sam Altman alone have this power and their inaction says everything we need to know about their intent.
Possibly, but not necessarily in a better direction. And that's the whole problem here. They should all go watch Phantasia.
If you don't tackle the societal level problems, nobody will care about research breakthroughs.
But something tells me either they can't or they won't, so no trust will be built.
I just don't buy the AI labs approach to this stuff. Like, unless we can basically simulate the entirety of human biology, I don't really see how LLMs can make progress here. Maths is different as it doesn't require a real-world interface, and programming already (by definition) can be simulated on a computer.
Without that, I can't see much (if any) progress being made on domains like biology.
Drug discovery is similar AFAIK. The space of possibilities is even larger than protein folding, but it's structurally similar enough that I think AI will help to make progress on the discovery side. Actually getting the drug tested and approved is another matter though for sure.
For one, the question for Anthropic is whether LLMs, specifically, not AI techniques more generally, can help significantly with cancer research. And here, all experience so far is that LLMs only really work when they can easily automatically verify their own outputs and self correct - such as in math (using automatic proof verifiers) or programming (using compilers and unit tests).
The second problem is that biological research speed is highly dependent on slow biological processes, such as cultures and long term studies. In programming, if an LLM could provide excellent insights and research suggestions 100x faster than a human, it would speed up the work roughly 100x. But in biology, it would only speed up the total work by a small amount - as any insight, even if absolutely brilliant and spot on, would still require months and years of actual experimentation.
I do agree that this kind of targeted approach makes sense.
However, discovery is not really the issue here. Running the clinical trials (1/2/3) is much much more difficult, and consumes basically all of the time in drug development, so even if LLMs perfectly automate this, the speedup will not be particularly large.
Anthropic is literally creating the bubble. It is not beyond their scope of influence, it is literally what they are consciously achieving.
As for peoples jobs, same actually applies. Anthropic is selling itself on dream of replacing jobs, even or especially where they are well aware AI does not perform that well. They are actively trying to replace people quickly before management notices it does not work well.
And also, they can influence how much their data centers contribute to global warming.
The way he's coming off here though gives me SBF vibes.
But instead of doing that, I'm going to drop everything and become a linux kernel maintainer. Learning about the internals of how RCU concurrency is implemented across subsystems is surely what I need to do right now!
(/j before I get crucified)
So Anthropic are going to spin up a medicinal chemistry lab and start mouse experiments?
I mean, it would be nice, but I don't think that the guys at Anthropic know what they're talking about here, or what they might be getting themselves into. (If indeed this is more than just PR.)
https://www.writingruxandrabio.com/p/intelligence-is-not-the...
Didn't and don't mean to disparage anyone's medical struggles, but "a cure for cancer" is a well-worn strawman. One that's been achieved for the low hanging fruit, the higher ones are seeing steady progress (already before LLM chatbots, even!), the bottleneck isn't "intelligence" and to the extent there are socioeconomic (access to screening, treatment) or environmental/lifestyle factors involved the AI boom is likely just making things worse!
I'm sure Dario knows this, and it's anyway too pedestrian compared to the usual list of fruits of ASI. The text probably originally read "nanobots eating you alive and uploading to the cloud" or something, but they figured that wouldn't go over with the intended audience. "What do the peasants care about? Oh I know! Curing cancer!"
How else are you going to work towards a dream? It's how Elon Musk managed to get reusable rockets when everyone said it's unfeasible.
Granted, most ideas don't work, this is why you need testing, but I am a bit surprised to see this attitude on hacker news.
But they said that they want to cure cancer. That's far outside the core competencies of any LLM, and it requires a lot of real-world wetwork with liquids, chemicals, cell line experiments, animal experiments, etc. They can't merely analyze existing data -- they'd need to generate vast amounts of new data, which isn't really the case in physics. And then regulatory approvals and so forth.
"We're working on curing cancer" sounds more like a poor PR attempt than an actual effort, though I'd love to be wrong.
“The detection solution may be made available in the Union as one or more of the following: (i) a public, ideally standardised, specification allowing any third party to implement a detection mechanism; (ii) a piece of software (e.g., a standalone executable or library); (iii) a cloud-based service accessible to users in the Union through an API.”
Option 3: "a cloud-based service".
And accessible "to users". NOT to the public.
Note: this is not the actual law, it is the code of practice, ie. guidance for model providers. The start of that document clearly states that you can comply with the law in other ways if you want, you'll just have to justify yourself. So you have a choice to not even do this.
The actual law is here:
https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5...
As to who decides the law clearly states who decides if someone is legal, like in most EU legislation. It's not the courts, it's "National market surveillance authorities designated by each EU Member State", and there is an EU office as well (and it is explicitly stated that they are not allowed to override each other). So every EU country has the right to provide exceptions to the law, just like they do for the GPDR. You do not have any rights under this legislation as an individual. Only these "National market surveillance authorities" get rights under this legislation.
Secondary: article 7 additionally gives the EU commission the power to declare any code of conduct they want that declares what compliance with the AI act actually means.
(and, of course, this is yet another attempt at declaring math illegal. The only way to actually enforce this legislation is for all models to comply with this, all over the internet. Obviously the EU does not remotely have the power to make that happen)
Where do I find that in the Act (or code of practice etc.)? IANAL, but Article 50 reads different to me but if there is a comment or guide how to read it - also fair enough.
> puts in it's regulations that you're not allowed to use AI to communicate with them (we all know this is coming),
.. do we?
> You get a mail from the government about taxes. You want to check if that text is generated by an AI
No, actually, I want to know whether it's correct and whether it's legally binding. Using AI makes it less likely to be correct, sure, but ultimately whether it's human, spreadsheet, or AI the important thing is the legal right to correct process.
Which is mostly a good thing because most of those processes were cooked up to be exclusionary but with plausible deniability
God forbid some upstart get a big contract they "shouldn't" have gotten because AI lets them create a submission that's on par with a big guys.
God forbid Joe Schmo be able to prepare the paperwork and dot his I's and cross his T's and be able to take advantage of some stupid zoning exemption that you spec'd out to essentially be only available to big moneyed interests.
The ability to suffer one's way through beurocratic process without having to suffer through paying the lawyers and accountants and engineers and whatnot to get you through that process is a huge part of the value proposition of AI. But people don't frame it like for obvious reasons.
Anthropic in particular has developed this almost Orwellian like veil of condescending rhetoric that on the surface suggests they're looking out for you while underneath they're taking actions that suggest they do not trust you, Mr/Mrs Ordinary Person. All the safety rhetoric, never supporting open weight models, the lockdowns on harnesses outside claude code, etc.
Anthropic if you care about public good, do something to empower people. Release an OSS model. Open source Claude Code. Open source some inference tooling or something. Just give people anything except your words.
Is this what kids call rage bait these days?
If you provide tools that help an organization kill easier, and possibly lazily rely on your inaccurate tools' judgment, and you do not care about the outcomes of these tools' uses, that's on your soul, too.
Yes, causality and attribution are hard sometimes, but it may have been that without it, hundreds of humans would be alive right now, and he doesn't even know. Maybe he should talk to the kids' parents about what alignment means.
This is actually a good point. Open models are getting better and better, some of them might even be useful in consumer hardware now. But if AI performance is still correlated with compute power, then no doubt power will remain with the people owning the chips.
The point "AI is structurally a technology that tends to concentrate power" is not that correct. They need this statement to be true, otherwise no way to justify the trillion evaluations.
Iran has hit the Amazon data center in Bahrain. The Ukranian deep drone bombing campaign has hit refineries, but also Wildberries, the "Russian Amazon" warehouses. The US campaign against Iran now, Iraq and Serbia previously, targeted power infrastructure. It will obviously be a target in the next war, and AI goes on that list too.
Every time the US completes an AI data center, someone in the Chinese nuclear command updates their target priority list. And vice versa.
You might have stumbled on the way to convince people that more data centers are a good thing
How is it not true?
The biggest companies in the world are, quite obviously (just look at the numbers) going to be AI companies. The most powerful governments in the world are quite clearly going to be those that are close to (or in control of) AI companies. People have a very hard time seeing the second order effects of the control of intelligence - we need to fix that.
The outputs of models in AI datacenters are the work.
The electricity company does not get the entirety of my useful input to the work as a result of me using electricity, but an AI company does.
Anthropic clearly are not though, and to call their operations wasteful is an understatement.
Are you self-hosting Google or Bing? No, but we have quite a huge ecosystem of full-text search tools with PageRank, with options to scale to almost Google scale (if you have the money). After all LLM training starts with the same crawl mechanism.
As long as barriers to entry is not too high (ie. it makes sense to take the risk to start a business that provides something similar - usually for a niche) market forces work.
We have the classic empirical chart reproducing microeconomics.
https://www.fda.gov/about-fda/center-drug-evaluation-and-res...
And setting up a pharma plant is also very capital intensive.
Here the obvious barrier to entry is completely artificial. (Which provides an incentive to spend a lot of money on R&D -- though it naturally raises the question of Pareto efficiency.)
1) in things like tax law, registering with city hall, dealings with the DMV, your phone subscription, insurance contract, ... you will find that one of the new fine prints in the contract will be that you're not allowed to use AI to communicate with them.
2) because of how SynthID works (you need the SynthID keys to verify, which are secret. So the only way to find if text is ChatGPT/Google/Anthropic watermarked is to ask ChatGPT/Google/Anthropic), government and large companies can enforce this against you. That is what the watermark is for. To end any insurance claim written by AI with "you're not allowed to submit AI written insurance claims" and refuse it outright there and then.
"Sorry your request was AI watermarked and pursuant to law 234 of 2025/03/11 chapter 3258 paragraph 33 decile 1299 we hereby close it without response"
3) when they reply, however, they use a custom model that also has custom SynthID keys. You will not even be able to tell their responses are AI written, or at least, you won't be able to prove it. You won't be able to enforce any AI-related rights (ie. the right to talk to a human) you have under the law against large companies.
In other words: this is to make sure that all the advantages AI provides are available to deny your unemployment claim, and to Verizon to charge you more, but completely inaccessible TO YOU when you want to change to a cheaper subscription. They can inundate YOU with AI-written requests BUT YOU CAN'T.
Self-hosting helps because it prevents them from verifying if your responses are AI written, because you can generate non-watermarked AI text and so there is a level playing field.
2) What actual law/regulation would that currently be that would be used for such an outright refusal?
3) Would that comply with current regulation?
It'll be spec'd out so that it's cheaper to bend over and take it than take it to court and prove them wrong.
For government, because it's in law or regulations (ministerial decisions in Europe). For large companies "You agreed to it" (you know, like you agreed to allow Verizon to sell your location data to Palantir)
The other 2 questions I don't understand. My point is that the EU AI directive makes this possible. Makes it possible in ONE direction, while prohibiting the other. AI can be used by government and large companies to spam you and deal with you, and can't be used by you without being 100% up front about that to them (ie. enabling refusal)
Just agreeing to it isn't necessarily enough at least in some countries.
2) What laws and sections specifically makes that possible? Are there examples of that happening?
3) Where can I find that interpretation of article 50(?)? Some other article? (To the extent that things would need to be labelled/watermarked etc.)
No. Here is the list of organizations that have the power to make laws in the EU (and JUST the across-the-EU part of that list, within countries, within states, within provinces, within towns there's another list). This is referred to in legal tradition as the "Hierarchy of norms", because there is also a clear order defined.
https://eur-lex.europa.eu/EN/legal-content/glossary/eu-hiera...
however, AI doesn't really influence this. already there's a lot of problem with things like Ticketmaster, Apple's walled garden, abuses of IP law (patent trolls, DMCA trolls), etc.
the insurance industry is a prime example of this. the suffering caused by power imbalance is incomprehensible, and yet there's not enough political will to address this.
sure, it's easily possible that some important aspects of our everyday lives will be worsened by bad AI regulation. but IMHO this is wholly an upstream problem, it's a symptom of bad politics. (a byproduct of the Zip2 to Tesla to "democracy with roman salute characteristics" pipeline.)
that said, obviously the foundation to have any chance of a nonpatological market to exist is that self-hosting has to be legal.
The #1 post on HN right now[1] is full of people jubilating about how they can run Qwen 3.8 27B on their > 5 year old GPUs. If that isn't democratization of AI, I don't know what is.
I'm sure he's smart enough to instantaneously realize this too, but as the famous Upton Sinclair quote goes, he won't mention it even if he does.
I think access to compute will matter just as much, if not more, as access to models.
> Even if we assume that models will no longer improve and we reach a point where everyone can run Fable in their laptop, surely running 1000x Fable agents would give you an advantage.
IMO, asymptotic advantages are marginal. At least for coding, we got a glimpse into how much of an advantage it gives (or doesn't) when the Claude Code codebase leaked[1] ~4 months ago. :)
If you want to leave it running with the fans going crazy for 40 mins or overnight or something, fair enough, but otherwise it doesn't seem worth it to me. It's certainly not "interactive", even taking into account the over-thinking it does by default.
The 3.6 (maybe they'll release a 3.8?) MoE model is much more usable (but obviously not as good) on this machine spec.
I'm not sure whether that's low or high, or how it compares to a general audience.
Making "tech" a career and a societal goal onto itself, without the adjoining understanding of and deep commitment to ethics and the responsible use of power, is why we're sliding into authoritarian rule by a small circle of techno-oligarchs.
We need less "move fast and break things" and more "plant trees you will not live to see bear fruit."
I'm an AI advocate but that question makes me feel that AI is simply "very useful" rather than being a historical game changer for humanity.
The earliest used models in this category would easily 10x (what did I just say) the creation of one-off short scripts, but you're not writing a 30-year app out of just a bunch of short scripts.
The METR time horizons graph suggests we can now get 10x (ahem) speedup on solving most coding problems that take us a few hours and about half of problems that take 2 days. Amdahl's law bites: even infinity speedup on half your problems is only 2x overall.
If you let an LLM loose, with a huge budget, what's the biggest artefact it can make before it drowns under the weight of bad decisions? The C compiler and web browser headlines a while back? The maths papers we see now that solve problems which stumped the maths world for decades?
But this is the other side of the same coin: teamwork. One person getting a thing made in twelve months vs a team of a hundred, you can scale up fast with money when you have a proof of concept, and an LLM can make a lot of proofs of a lot of concepts even in free accounts.
I do agree that the port would've taken a lot longer without LLMs though.
After an aquisition earlier this year I got the task of doing an SAP-Integration for the new company, last time I did this 5 years ago it was a 6 month task, but with the experience and skills ive gained since I estimated it would be a 3 month project (with or without AI, most work is just logistics, AI cant help much there).
In those 3 months I was able to not only integrate SAP but also deliver a completely modernised user-facing software for that integration. While I could have written that software myself in a vacuum it would have never been worth it financially, since it would have delayed the launch of the integration by 6+ months. Building the software post-launch of the integration would have easily taken 2.5 years at minimum.
But this is also basically a "spherical cow in a vacuum" scenario, where I was essentially acting as a solo dev, in full operational control of the project, with deep domain knowledge of the topic and an allready fully set up codebase that I knew perfectly while working down ideas I've had in my backlog for 5+ years.
What you have said is correct, it lets you build software much faster. The question however is: is that software making money for the company? (Not talking about what you built but in general)
I think, with AI, companies are saying yes to a lot of things they would have said No to ik say 2020. And as a result realizing “just building it” is not the answer.
Previously your GTM team or Product team would say “If we ship some big project X, we unlock $Y in revenue” but now people are realizing that those projections were really more of a hope. So companies are spending so much more tokens and shipping so many more PRs based on hope but a lot of it just doesn’t turn into meaningful revenue, especially not in short term
This sort of system only works with internal software and an unusual amount of data. If we were in the business of selling that software we could not have charged a higher price for the new version over the old, the tweak could only be unlocked because we were able to control staff hiring and staff onboarding fully to make use of the new changes.
As part of the aquisition I got access to their previous codebase which was some sort of incomprehensible PHP monolith, with the persons who wrote that code long gone. Thanks to LLMs I was actually able to extract the core useful concepts (again, sufficiently deep domain knowledge that I knew exactly what to look for). Without LLMs i would have probably extracted the absolute minimum and let the rest rot.
There is no reason a dev of comparable skill and domain knowledge would not be able to do that for what I built here.
So yes, many aspects of my job are now 10x as productive, but turns out that improves my overall throughput only very little.
Second, the Internet didn't show up much in GDP and similar measures either!
But your point stands. Where are the amazing digital products/stuff? I get that it might take time to arrive as we scale up compute and learn new paradigms. But so much infra already exists (deployment pipeliens, everyone reachable on a smartphone) that we should be seeing something.
I'm definitely seeing indie-sized games that appear to have had significant input from AI, though I'm not sure the balance between AI for coding and AI for assets. My experience attempting this directly suggests that the current level they work at can make very simple games as one-shots, but anything more than trivial will produce outputs only as good as the developer's combined willingness to put in effort tweaking things and taking it all one step at a time, and their taste about what "good" even is.
I'm using spare credits to build and improve an isochrone map renderer, which I otherwise wouldn't have had time for (apart from anything else, I'd have had to become skilled in JS+wasm, somewhat of a pivot from iOS). This also requires taking it all one step at a time, having UX and UI taste.
Having lived through GeoCities since before it was bought by Yahoo!, taste is… well. Most people make things that nobody else actually wants.
But I am just as amazed with how little real life consequence it seems to have! Even software houses were hit more by interest rates than by this magical revolution.
If I couldn't directly observe Fable in action, I wouldn't believe in AI.
>> What's moving the goalposts? I am very much amazed at what Opus 4.8 can do. I push its code straight to prod.
>>> What's moving the goalposts? I am very much amazed at what Opus 4.6 can do. I push its code straight to prod.
>>>> What's moving the goalposts? I am very much amazed at what GPT5 can do. I push its code straight to prod.
>>>>>> What's moving the goalposts? I am very much amazed at what Opus 3.5 can do. I push its code straight to prod.
I've done amounts of refactoring and fixes and written tooling that just wouldn't have happened before.
I'm not sure what amazing new stuff y'all expect but the amount of technical debt in my projects is actually going down, cause I can finally get good enough test coverage, including E2E/load tests that actually prove whether the software works and scales or doesn't - just last week I diagnosed issues with SeaweedFS failing under concurrent writes when backing Sentry and could swap it out for Garage in a day, caught by a monitoring tool I slopped together that integrates with the Sentry API, no issues since.
The environment around me has gone from drowning in tech/ops debt to sort of swimming and at least holding above water for now (cause nobody will pay for 5x more tokens).
It's also insanely good for prototyping and being able to actually explore various ideas and shoot the bad ones down quickly instead of handwaving and looking at a loaded calendar, alongside being able to address well bounded tasks in parallel, better than human developers can - like I can give 5 GitHub issues to the slop machine and have it fix all of the annoying bugs. Issue with how some data shows up? Just feed it the DB dump and let it find out what's up.
Some projects have gone from around 500 code tests to around 4000, and before anyone says they're meaningless, at least 5% of those have caught real issues and helped a bunch, alongside linters and other tooling (including some tools I wrote myself). I've also written both native utilities and some web platforms for myself, side projects that I never would have gotten around to.
I'm measurably more productive than I've ever been (since I did measure that, looking at my commits over the last 2 years) but also burnt out. Still, it's the kind of burnout that's the consequence of context switching and lots of work, rather than the kind that I had years ago, where I had to manually untangle deeply nested Spring Boot service logic all over the place at like 2 AM cause the made up deadlines were kicking my butt.
In contrast to others, I don't need to move the goalposts - the productivity for me is here and now. Any future models will just make it better, unless we experience model collapse.
Disclaimer: you do need a LOT of code tests and validations, otherwise it all goes to shit. Maybe I'm just extending how much time it will be until it goes to shit for me as well, but go figure. You also have to babysit the models more than anyone would like or should, most of my work usually has 20-60 minutes of planning before dispatching the agent.
The way it changes the game is by lowering the cost of making radical bets so we end up trying more moonshots.
Or an example of MS - their main cost like most software companies are people, especially software devs, which are to be replaced by AI so on the surface they would greatly benefit from it. But their products are centered around helping out people do stuff on the computer. Why would you need that when the AI will do it better and faster directly operating on the data or using e.g. Python?
A product still requires a lot of handholding and human thinking, at least if one does not want everyone even throwing a glance at it to immediately be repulsed by the usual AI slop tells.
> I'm an AI advocate but that question makes me feel that AI is simply "very useful" rather than being a historical game changer for humanity.
It absolutely already is a historical game changer on par with the Industrial Revolution when it comes to the amount of jobs destroyed and economies screwed up - and the impact will be even worse in 10+ years as existing seniors retire but no new seniors rise as AI has destroyed entry level career paths.
Which economies are already screwed up?
IMO it can’t ever be on par with the Industrial Revolution because AI can only really affect the information economy. Things people do with their hands/bodies have either already been automated or can’t be with current tech. If you’d asked people decades ago they might say no one will ever work in factories by 2026 because they’ll all be automated. It didn’t work out that way. I think AI will go the same way: absolutely game changing to some industries (of which software engineering will be one) but a great many will still survive with less dramatic changes.
If anything it might result in more focus on the human aspects. How many people out there earn their stripes putting together slide decks? In a world where an AI can put together the snazziest presentation you’ve ever seen in a heartbeat it’s going to matter more how you stand at the front of the room and present those slides than it does today.
Law abiding US citizens won't be able to run them, but you didn't solve anything other than making sure US citizens pay Sam or Dario.
Or are we still doing the lawfare strat since you can’t get any meaningful lead on Chinese models?
How about the involvement in the Iran bombing that killed 160, out of which 120 were kids?
Forgive me if Anthropic doesn’t really rhyme with trust over here.
Make it "Apologies for linking to Amodei" and we're talking.
Or even for the HN crowd, when will we see e.g. "Mythos aided research discovers 10 new viable battery technologies"
> Open-weights do help some with this but are nowhere near a sufficient solution
Which is exactly why Anthropic contributes nothing to, and actively pushes for roadblocks and regulations for open weight models. Can't allow any hope to the masses.
Only people with $$$$$ are allowed to touch Fable. Which of course we have have aggressive guardrails for in case you even try to use it for something dangerous like AI model developement. Plus we are going to retain all data submitted to it just in case someone is trying to be sneaky.
this guy need to stop vibeposting nonsense.