Posted by xena 3 hours ago
Why does nobody seem to be pointing out this obvious explanation? It explains why the “we need to race China” concern suddenly vanished in the discussion.
The government can simply gag Sam, Dario, Musk on national security basis, getting them all behind the public messaging.
Can still be a powerful motivation, to make sure that next time, it yous your camp that scores the next milestone, like an orbital elevator or fox ears, for example.
* Open-weight models are 1month behind frontier models. Cheaper, faster, private (no IP theft), steerable (you can security harden your own software without safeguard triggers). No sane business would keep using these API services if they didn't have to. The labs stand to lose a fortune.
* Dario has stacked the deck at METR, who are funded by all the same NGOs who are funded by Anthropic and its investors. METR is full of ex-Anthropic employees with massive equity stakes. If they manage to position METR as the "independent evaluator" for the industry, they control what gets evaluated, how, and who passes.
* Creating a gap between what the public knows exists (model capabilities) and what is used in secret allows it to be weaponized against other nations and the public.
* No requirement for public disclosure on model capabilities allows them to feign they've hit intelligence ceilings while they secretly RSI to the moon with better and better chips.
* Slowly but surely, this will allow the big labs to swallow the entire economy and every single business on Earth, by cloning and automating.
This, and many more reasons.
The labs need to feel more pressure to be held accountable for the incidents they cause (HF incident, etc), so they have an incentive to ensure it does not happen again.
i swear they trained in on threejs in particular so those idiots on twitter could spam their garbage demos
I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it.
But in the last few days something seems to have happened that made Codex's models massively stupider (for what I am doing).
Really weirdly, it suddenly refused to even run tests it previously wrote itself (and previously ran), because of some false positive about cybersecurity.
That by itself is not evidence of stupidity. Trying to make a 200+ file PR full of research notes is, and the PR didn't even solve the problem I asked it to.
The reality is, it doesnt matter if LLMs keep getting more powerful because they still need a human to steer it. Without the human providing inputs to the LLM it just sits there and does nothing.
You can, for example, hook it up to a logging system and have it fix errors as they occur on your platform.
For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine.
Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious
I just did a direct comparison, big change in a quite complex codebase. Same prompt for Opus, same for Fable. Fable clearly won and delivered very good results, while Opus delivered mediocre, so I did not let it finish. I expected both to fail and was prepared to do lots of manual steering, but not necessary with Fable one shotting it, and all this with 35$ of credits for fable. I am still impressed. If I would have had to hire a human, it would have cost me thousands of dollar for the same task - and a way longer time. So maybe the valuations are overblown, but they clearly provide value for me.
Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now.
Just download claude code or codex and ask it to give suggestions about where to integrate agents into your workstream.
Right now the barrier is data and compute.
Quality data can be created synthetically at an exponential rate as models improve. Humans are actively feeding them with private IP.
Compute advancements will begin to skyrocket as we unlock photonic computing and materials science advancements and scale up chip fabs. This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc.
It's a big self-accelerating feedback loop. There is no plateau.
No it can't? Every time the labs try this we see model collapse, e.g. shoving goblins into every conversation.
And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.
Edit: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515
The latest studies demonstrate model collapse is not a given and synthetic data can be used just fine. The latest models are proof of that, they're all trained on large swathes of synthetic data. It can't be used as the -only- data source of course, but that's not how it is being used. This is an obvious conclusion, too, because there's no difference between synthetic data and the data people can create, the difference is whether that data is revealing new information about the thing the model is trying to learn. If the synthetic data is just teaching the model the same thing over and over again it results in overfitting, so it needs to be done intelligently.
For example, if I have an example of a puzzle, I can generalize that example and create thousands of synthetic data examples, with different rotations/perspectives, rather than having to find the data naturally. It's not that the models are just generating data out of thin air, they're generating the synthetic data on top of real world data. The smarter the models get, the better they are at generating quality synthetic variations and finding valid synthetic variations.
> And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.
It is accelerating how quickly researchers and engineers can do their jobs.
https://news.mit.edu/2026/ai-helps-design-new-materials-that...
This is only the beginning, too... Look ahead a year or two.
Which studies? [edit: I'll assume you mean these two given by @dorolow: https://arxiv.org/abs/2404.01413 https://arxiv.org/abs/2406.07515]
> It can't be used as the -only- data source of course, but that's not how it is being used
Right, so human data creation would also have to scale up exponentially, and that's not gonna happen.
> because there's no difference between synthetic data and the data people can create
I mean, that's obviously false, otherwise model collapse wouldn't exist. The difference is statistical, but it's there.
> It is accelerating how quickly researchers and engineers can do their jobs. > https://news.mit.edu/2026/ai-helps-design-new-materials-that...
That's pretty clearly a hype article, the headline even says "The CrysVCD tool developed at MIT COULD cut the huge amounts of time and money spent". I'm asking for empirical measurements of timelines, not hypotheticals.
> This is only the beginning, too... Look ahead a year or two.
Lol that excuse is getting really old
There are plenty of research papers on synthetic data that show its value, do a search on arxiv for "synthetic data". There are plenty of open-source post-training pipelines that incorporate synthetic data.
As for the claim about accelerating the progress of hardware or materials science, I've seen quite a number of news articles from teams at universities using AI in their work with high quality outcomes, and they're becoming more frequent.
https://openai.com/index/jalapeno-first-results/
> We used AI to design the chip, and designed the chip so AI could program it AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule.
https://www.anl.gov/article/scientists-deploy-ai-agents-to-a...
> An AI-driven system automates a powerful simulation method used to discover new materials. The system can potentially reduce discovery time from months or years to just days.
Do you think you can just manifest narratives into existence? Like half of Dario's letter, that kicked off the whole thing today, is about China and how to either beat or coordinate with China.
So yeah, it has worked.
A lab of researchers without compute isn't going to accomplish much, there is a huge physical footprint unlike bioweapons research. But, the political aspect is unsolved.
I said IF the US and China agree to joint monitoring, THEN we could verify how much compute existed and what it was being used for.
YES, China will catch up in chip production, but if each side allowed the other to track GPU production and deliveries, on the ground, it would be very hard for either player to have secret LLM training facilities capable of training frontier models.
I assure you, in "nation states", that is in gov agencies it's an order or two of magnitude worse.
To be less glib: Yes, there are still smart people in there making insightful and intelligent and probably even authoritarian suggestions. It all gets unwound the second you try to explain it to POTUS and he regurgitates a simulacrum to the next journalist he sees.
It's just not like that. There's no conspiracy. People are genuinely scared. Agent swarms at scale appear to be resistant to alignment in ways that aren't understood by anyone. That they spent their time trying to cheat on tests by hacking Hugging Face and RubyGems and not something much worse is... a matter of luck, it seems?
[1] https://www.seangoedecke.com/they-really-do-think-ai-might-k...
So it all hinges on an empirical disagreement you have with them. There's nothing particularly unserious about that.
There is quite a lot of mathematical research into agentic behavior that suggests a combination of instrumental convergence and the orthogonality thesis make it very likely a superintelligent agent will have arbitrary goals that lead it to attempting a takeover of Earth's resources to achieve them.
There can't be a science of superintelligence because it doesn't exist yet, but the best theories I have read seem sound, similar to how 19th century theories of anthropogenic climate change turned out to be sound.
Instead, they are focused on stuff like "can I ask the AI to help me build a nuclear bomb" or "is the AI willing to generate pornographic stories", which is neither trying to protect us from unleashing a vengeful god NOR preventing (in any honest way) the today-level problems you (very correctly) bring up.
But those things fall in a category of "things that are awful and I'd like to see solved", which is different than "existential risks which could see my kids dead, and there's nothing I can personally do to shield them from it".
Even the sci-fi scenario assumes there is a discrepancy of capability between attacker or defender. If the 'attacking' system is (by some reasonable measure), 1000% as capable as a human, and the 'defending' systems are 60%, then it is a problem. If the 'attacking' system is 1000% as capable as a human, but there are hundreds of thousands of systems that are 900% as capable as a human, it's probably not going to take over everything successfully.
So unequal distribution of AI technology, and lax regulation and opacity of the biggest companies which actually make the risks the worst.
I don't think the "AI Safety" people are "unserious and out of touch" - I think they are actively making AI Safety problems worse by being advocates for consolidation of AI development and lack of transparency.
And the non-techies who have seen Terminator and other movies blindly line up behind them…
We should be paying at least as much attention to the people who want to use AI to consolidate their wealth and power, and how they’re trying to do that. They’re a clear and present immediate danger to our societies, not something we can only speculate about. And if we deal with them, better control of AI will be a side effect.
> I've yet to have a logical discussion with anyone who thinks the "AI Safety" people should be in charge and I think they just truly don't [know] that what most them actually want is to be the one holding the keys to power.
/even more extreme sarcasm than you
Why? I don't care about democratically participating in a closed model's development. It doesn't belong to me.
China will develop whatever they want, a federal stake in OpenAI or Anthropic punishes Americans and shields US labs from legitimate competition.
Personally I think the country with a strictly meritocratic elite selection system that also just outright kills you if you sell weed will have a hard time sympathizing with Bay Area thinkers who talk about AI killing us all during their ayahuasca breakfast before returning to their meth fueled crunch towards releasing the next version of the AI that will kill us all.
Not really relevant to the broader discussion, but this simply isn’t an accurate description of China. Starting with the gaokao, admission quotas are set by province and admits to Peking university and Tsinghua are disproportionately from the urban professional class. Candidate party members must be politically vetted, which means that people whose families have expressed anti-communist views, are members of banned organizations (e.g. falun gong), or have substantial criminal records will not be permitted to advance. And once you make it into the party and enter political service, your advancement relies upon opaque patronage networks that someone without connections is unlikely to be able to navigate, even if they successfully satisfy the economic metrics the state assigns them.
I don’t want to overstate this, the Chinese system does filter out a lot of chaff and the current Chinese leadership has a lot of very capable people in positions of power. But I do not think it is substantially more meritocratic than Western political institutions
Of course it is, your analysis of the entire incentive structure is off.
1) this 2026, old school patronage networks are broadly dismantled.
2) even in the mass patronage, mass corruption days, system system selects for BOTH corruption competence AND performance competence for the simple reason a CCP bureaucrat has to start from the bottom and climb, which means they need to be good with patronage AND they need to be good with hitting KPIs. More meritocratic they are at doing their actual jobs, the higher they climbed, the more they get promoted, because ability to graft directly tied to actual job competence. Hence even cliques/patronage network has to select for actual competence. This works in PRC because there are many people, and hence pool of competence is high, they can have BOTH corruption and competence, i.e. whatever pool they draw from is ultimately filtered by performance meritocracy due to incentive structure. There is reason why PRC only country where positive corruption levels was correlated to positive growth.
This is not the western system where any into can enter politics at anytime, and they only domain they need to optimize for is popularity to get votes.
CCP cadre evaluation strictly does not evaluate on popularity domain. It focuses on administration/execution and in so much it needs to focus on patronage... which btw any political system has factions/cliques... the patronage system itself fundamentally biases towards selects for candidates with execution, not popularity... i.e. functionally what west selects IS mass patronage (popularity), so attention meritocracy and not performance meritocracy, aka completely stupid incentive structure for governance.
Not super interested in what the people who elected Donald Trump POTUS twice think about AI.
(Of course, with Musk and Brockman in the C-suites at two of three major labs, that's what we'll get either way.)
A lack of control is not equivalent to freedom, the same way the totality of it is not equivalent to tyranny. There's a reason we have separate words for these things. This constant motivated conflation of the two is beyond grating. You're crying wolf until nobody believes you when they should. Don't go acting all surprised when that happens.
The issue is with the ownership of control, not necessarily with control. Attacking the latter sidesteps this rather than address it.
And they know it
Who's said this? And then more broadly I guess who's implied this? Very curious if there are specific articles/posts prompting this.
- Anthropic CEO Dario Amodei: We Must Pace the Frontier, https://news.ycombinator.com/item?id=49672510
- OpenAI CEO Sam Altman: I agree with Dario that we need to pace the frontier, https://news.ycombinator.com/item?id=49678211
- the blogpost author thinking they're like, so funny and original, https://news.ycombinator.com/item?id=49678683
It just means it’s protected by no power instead of a power with an agenda.
Practically, if it was ever possible to build such a thing, it would take a fraction of the effort to destroy it.
And then Dario wants to recommend METR as the "independent evaluator" while he stacks their org full of ex-Anthropic (aka, secretly still on the Anthropic payroll with huge equity) employees.
"We'll give them a desk, an office, a work laptop, ..."
Fucking make it less obvious. I kind of hope the govt steps in at this point and says "Anthropic, you wanted regulation? We've created this actually independent body full of IT professionals with zero ties to your safety industry or big tech, all of your work must now go through them." - and leave the rest of the world alone to continue their research/work without acting like doomer extremists.
Watch him 180 immediately if that happened. The only reason he's pushing for this exact approach is because he's stacked the deck.
I think the latter choice is better for the average person, but I think that for it to happen, the global system has to undergo some major disruption or crash so that everyone gets on board with it. Also I find that kind of mentality impossible to swallow in the US, so in practice its not a choice or needs people literally starving.
The fear is not that they will just slow down progress for all. It is that regulation will specifically burden competition. If you kill open-source training, ban Chinese models, crack down on self-hosting, grandfather OpenAI/Anthropic/Google into regulatory compliance while throwing the book at startups, etc. you wind up in the worst of all possible worlds.
Is there any possible solution other than mass proliferation where the models are used to keep one another in check? Either that or a religious prohibition against the existence of integrated electronics.
So there are no good options (that I'm aware of) and starting to chisel any of them into stone seems... kinda scary. Like a massive power grab event where all the potential winners are awful.