Posted by Philpax 2 hours ago
Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...
one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?
> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.
Early access, no weights no tech details, just a sign up here for info
Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.
This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.
> We will release the weights, technical report, model card, and developer artifacts later this month.
Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.
Am I missing something?
It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.
I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models
> That only means their training regime is inferior if their predecessors did so much more with so much less
Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.
None of that means they shouldn’t release their model.
500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.
seems like they are aiming to provide both inference and RLaaS for american companies and western govts. even if they never fully beat deepseek if they get close enough the fact that they're American will help them close deals
It's great to see a company that acknowledges it still needs improvement instead of making false claims.